vienalatina/docker/translate/Dockerfile
Claude 24dbc9b961
Switch translation to OPUS-MT; mount models instead of baking them
M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").

The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.

OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.

Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:

  es -> de       opus-mt-es-de           (small, ~74M)
  de -> es       opus-mt-tc-big-de-es    (tc-big, ~237M)
  es <-> pt-br   opus-mt-tc-big-itc-itc  (>>pob<< / >>spa<<)
  de <-> pt-br   no model in either direction — pivots through Spanish

Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.

MT_PROVIDER=m2m100 still selects the old engine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
2026-09-18 12:41:45 +00:00

34 lines
1.1 KiB
Docker

# Pipeline image for the translate step.
#
# Models are NOT baked in — they live on the host at /srv/mt-models (see
# scripts/fetch-models.sh) and are mounted read-only by .woodpecker.yml. That
# keeps this image small and lets a model change be a directory swap instead of
# a 1GB image rebuild, which matters because model choice turned out to need
# iteration.
#
# torch is deliberately absent: it is only needed to *convert* a model, and the
# default PyPI wheel drags in ~2.5GB of CUDA libraries this CPU-only box will
# never use. Conversion happens in fetch-models.sh, not here.
#
# Build on the server:
# docker build -t vienalatina/translate:2 docker/translate
FROM python:3.12-slim
# translate.py shells out to git to diff the push and commit the siblings back.
RUN apt-get update -qq \
&& apt-get install -qq -y --no-install-recommends git \
&& rm -rf /var/lib/apt/lists/*
RUN pip install --no-cache-dir \
ctranslate2 \
transformers \
sentencepiece \
sentencex \
pyyaml
ENV MT_PROVIDER=opus \
MT_MODEL_DIR=/opt/mt/models \
MT_COMPUTE_TYPE=int8 \
MT_THREADS=2