vienalatina/scripts/translation/opus_provider.py
Claude 24dbc9b961
Switch translation to OPUS-MT; mount models instead of baking them
M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").

The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.

OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.

Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:

  es -> de       opus-mt-es-de           (small, ~74M)
  de -> es       opus-mt-tc-big-de-es    (tc-big, ~237M)
  es <-> pt-br   opus-mt-tc-big-itc-itc  (>>pob<< / >>spa<<)
  de <-> pt-br   no model in either direction — pivots through Spanish

Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.

MT_PROVIDER=m2m100 still selects the old engine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
2026-09-18 12:41:45 +00:00

84 lines
3.5 KiB
Python

"""OPUS-MT (Helsinki-NLP) via CTranslate2, on CPU.
Replaces M2M100 418M, which was measured producing unusable German on real
articles — "frijoles" became "Beeren" (berries), "rompe la idea" became
"breitet die Idee" (spreads it), and the site's own name came back mangled.
Not every pair exists as a published model, so routes are explicit rather than
assumed. Verified against Hugging Face:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (Italic multilingual)
de <-> pt-br no direct model — pivots through Spanish
Marian models take the target language as a `>>xxx<<` token at the start of
the *source* text, unlike M2M100's decoder-side target_prefix.
"""
from __future__ import annotations
import os
from .provider import Provider
# (source, target) -> [(model directory, target-language token or None), ...]
# More than one hop means a pivot translation.
ROUTES: dict[tuple[str, str], list[tuple[str, str | None]]] = {
("es", "de"): [("opus-mt-es-de", None)],
("de", "es"): [("opus-mt-tc-big-de-es", None)],
("es", "pt-br"): [("opus-mt-tc-big-itc-itc", ">>pob<<")],
("pt-br", "es"): [("opus-mt-tc-big-itc-itc", ">>spa<<")],
("de", "pt-br"): [("opus-mt-tc-big-de-es", None),
("opus-mt-tc-big-itc-itc", ">>pob<<")],
("pt-br", "de"): [("opus-mt-tc-big-itc-itc", ">>spa<<"),
("opus-mt-es-de", None)],
}
class OpusMTProvider(Provider):
def __init__(self, model_root: str | None = None, compute_type: str | None = None,
threads: int | None = None):
self.model_root = model_root or os.environ.get("MT_MODEL_DIR", "/opt/mt/models")
self.compute_type = compute_type or os.environ.get("MT_COMPUTE_TYPE", "int8")
self.threads = threads or int(os.environ.get("MT_THREADS", "2"))
self._name: str | None = None
self._translator = None
self._tokenizer = None
def _load(self, name: str):
"""Keep exactly one model resident — three at once would not fit a 4GB box."""
if self._name == name:
return self._translator, self._tokenizer
import ctranslate2
import transformers
self._translator = None # free the previous model before allocating the next
path = os.path.join(self.model_root, name)
self._tokenizer = transformers.AutoTokenizer.from_pretrained(path)
self._translator = ctranslate2.Translator(
path, device="cpu", compute_type=self.compute_type, intra_threads=self.threads
)
self._name = name
return self._translator, self._tokenizer
def _hop(self, texts: list[str], model: str, token: str | None) -> list[str]:
translator, tok = self._load(model)
prepared = [f"{token} {t}" if token else t for t in texts]
batch = [tok.convert_ids_to_tokens(tok.encode(t)) for t in prepared]
results = translator.translate_batch(batch)
return [
tok.decode(tok.convert_tokens_to_ids(r.hypotheses[0]), skip_special_tokens=True)
for r in results
]
def translate(self, texts: list[str], src: str, tgt: str) -> list[str]:
if not texts:
return []
route = ROUTES.get((src, tgt))
if route is None:
raise ValueError(f"no translation route from {src} to {tgt}")
for model, token in route:
texts = self._hop(texts, model, token)
return texts