M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").
The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.
OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.
Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (>>pob<< / >>spa<<)
de <-> pt-br no model in either direction — pivots through Spanish
Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.
MT_PROVIDER=m2m100 still selects the old engine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
34 lines
1.1 KiB
Docker
34 lines
1.1 KiB
Docker
# Pipeline image for the translate step.
|
|
#
|
|
# Models are NOT baked in — they live on the host at /srv/mt-models (see
|
|
# scripts/fetch-models.sh) and are mounted read-only by .woodpecker.yml. That
|
|
# keeps this image small and lets a model change be a directory swap instead of
|
|
# a 1GB image rebuild, which matters because model choice turned out to need
|
|
# iteration.
|
|
#
|
|
# torch is deliberately absent: it is only needed to *convert* a model, and the
|
|
# default PyPI wheel drags in ~2.5GB of CUDA libraries this CPU-only box will
|
|
# never use. Conversion happens in fetch-models.sh, not here.
|
|
#
|
|
# Build on the server:
|
|
# docker build -t vienalatina/translate:2 docker/translate
|
|
|
|
FROM python:3.12-slim
|
|
|
|
# translate.py shells out to git to diff the push and commit the siblings back.
|
|
RUN apt-get update -qq \
|
|
&& apt-get install -qq -y --no-install-recommends git \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
RUN pip install --no-cache-dir \
|
|
ctranslate2 \
|
|
transformers \
|
|
sentencepiece \
|
|
sentencex \
|
|
pyyaml
|
|
|
|
ENV MT_PROVIDER=opus \
|
|
MT_MODEL_DIR=/opt/mt/models \
|
|
MT_COMPUTE_TYPE=int8 \
|
|
MT_THREADS=2
|