M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").
The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.
OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.
Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (>>pob<< / >>spa<<)
de <-> pt-br no model in either direction — pivots through Spanish
Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.
MT_PROVIDER=m2m100 still selects the old engine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
47 lines
1.5 KiB
Bash
47 lines
1.5 KiB
Bash
#!/usr/bin/env bash
|
|
# Download and convert the OPUS-MT models the pipeline needs.
|
|
#
|
|
# Run once on the server; the pipeline mounts the result read-only, so changing
|
|
# models later means re-running this rather than rebuilding a 1GB image.
|
|
#
|
|
# bash scripts/fetch-models.sh # -> /srv/mt-models
|
|
# bash scripts/fetch-models.sh /tmp/mt # -> somewhere else
|
|
#
|
|
# Conversion holds the fp32 model in RAM. These are 74M-237M parameter models,
|
|
# so each peaks well under the 4GB box's headroom — unlike M2M100 418M, whose
|
|
# conversion was OOM-killed here at a 3GB cap. One container per model keeps a
|
|
# failure contained to that model.
|
|
|
|
set -euo pipefail
|
|
DEST="${1:-/srv/mt-models}"
|
|
|
|
MODELS=(
|
|
"Helsinki-NLP/opus-mt-es-de"
|
|
"Helsinki-NLP/opus-mt-tc-big-de-es"
|
|
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
|
|
)
|
|
|
|
sudo mkdir -p "$DEST"
|
|
sudo chown "$(id -u):$(id -g)" "$DEST"
|
|
|
|
for REPO in "${MODELS[@]}"; do
|
|
NAME="${REPO##*/}"
|
|
if [ -f "$DEST/$NAME/model.bin" ]; then
|
|
echo "== $NAME already converted, skipping"
|
|
continue
|
|
fi
|
|
echo "== converting $NAME"
|
|
docker run --rm -u "$(id -u):$(id -g)" -e HOME=/tmp --memory=3g \
|
|
-v "$DEST:/out" python:3.12-slim bash -c "
|
|
set -e
|
|
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
|
|
--extra-index-url https://download.pytorch.org/whl/cpu
|
|
ct2-transformers-converter --model $REPO --output_dir /out/$NAME \
|
|
--quantization int8 --copy_files source.spm target.spm vocab.json tokenizer_config.json
|
|
"
|
|
done
|
|
|
|
echo
|
|
du -sh "$DEST"/*
|
|
echo "Models ready in $DEST"
|