vienalatina/scripts/fetch-models.sh
Claude 24dbc9b961
Switch translation to OPUS-MT; mount models instead of baking them
M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").

The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.

OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.

Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:

  es -> de       opus-mt-es-de           (small, ~74M)
  de -> es       opus-mt-tc-big-de-es    (tc-big, ~237M)
  es <-> pt-br   opus-mt-tc-big-itc-itc  (>>pob<< / >>spa<<)
  de <-> pt-br   no model in either direction — pivots through Spanish

Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.

MT_PROVIDER=m2m100 still selects the old engine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
2026-09-18 12:41:45 +00:00

47 lines
1.5 KiB
Bash

#!/usr/bin/env bash
# Download and convert the OPUS-MT models the pipeline needs.
#
# Run once on the server; the pipeline mounts the result read-only, so changing
# models later means re-running this rather than rebuilding a 1GB image.
#
# bash scripts/fetch-models.sh # -> /srv/mt-models
# bash scripts/fetch-models.sh /tmp/mt # -> somewhere else
#
# Conversion holds the fp32 model in RAM. These are 74M-237M parameter models,
# so each peaks well under the 4GB box's headroom — unlike M2M100 418M, whose
# conversion was OOM-killed here at a 3GB cap. One container per model keeps a
# failure contained to that model.
set -euo pipefail
DEST="${1:-/srv/mt-models}"
MODELS=(
"Helsinki-NLP/opus-mt-es-de"
"Helsinki-NLP/opus-mt-tc-big-de-es"
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
)
sudo mkdir -p "$DEST"
sudo chown "$(id -u):$(id -g)" "$DEST"
for REPO in "${MODELS[@]}"; do
NAME="${REPO##*/}"
if [ -f "$DEST/$NAME/model.bin" ]; then
echo "== $NAME already converted, skipping"
continue
fi
echo "== converting $NAME"
docker run --rm -u "$(id -u):$(id -g)" -e HOME=/tmp --memory=3g \
-v "$DEST:/out" python:3.12-slim bash -c "
set -e
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
--extra-index-url https://download.pytorch.org/whl/cpu
ct2-transformers-converter --model $REPO --output_dir /out/$NAME \
--quantization int8 --copy_files source.spm target.spm vocab.json tokenizer_config.json
"
done
echo
du -sh "$DEST"/*
echo "Models ready in $DEST"