M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").
The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.
OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.
Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (>>pob<< / >>spa<<)
de <-> pt-br no model in either direction — pivots through Spanish
Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.
MT_PROVIDER=m2m100 still selects the old engine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
84 lines
3.5 KiB
Python
84 lines
3.5 KiB
Python
"""OPUS-MT (Helsinki-NLP) via CTranslate2, on CPU.
|
|
|
|
Replaces M2M100 418M, which was measured producing unusable German on real
|
|
articles — "frijoles" became "Beeren" (berries), "rompe la idea" became
|
|
"breitet die Idee" (spreads it), and the site's own name came back mangled.
|
|
|
|
Not every pair exists as a published model, so routes are explicit rather than
|
|
assumed. Verified against Hugging Face:
|
|
|
|
es -> de opus-mt-es-de (small, ~74M)
|
|
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
|
|
es <-> pt-br opus-mt-tc-big-itc-itc (Italic multilingual)
|
|
de <-> pt-br no direct model — pivots through Spanish
|
|
|
|
Marian models take the target language as a `>>xxx<<` token at the start of
|
|
the *source* text, unlike M2M100's decoder-side target_prefix.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
|
|
from .provider import Provider
|
|
|
|
# (source, target) -> [(model directory, target-language token or None), ...]
|
|
# More than one hop means a pivot translation.
|
|
ROUTES: dict[tuple[str, str], list[tuple[str, str | None]]] = {
|
|
("es", "de"): [("opus-mt-es-de", None)],
|
|
("de", "es"): [("opus-mt-tc-big-de-es", None)],
|
|
("es", "pt-br"): [("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
|
("pt-br", "es"): [("opus-mt-tc-big-itc-itc", ">>spa<<")],
|
|
("de", "pt-br"): [("opus-mt-tc-big-de-es", None),
|
|
("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
|
("pt-br", "de"): [("opus-mt-tc-big-itc-itc", ">>spa<<"),
|
|
("opus-mt-es-de", None)],
|
|
}
|
|
|
|
|
|
class OpusMTProvider(Provider):
|
|
def __init__(self, model_root: str | None = None, compute_type: str | None = None,
|
|
threads: int | None = None):
|
|
self.model_root = model_root or os.environ.get("MT_MODEL_DIR", "/opt/mt/models")
|
|
self.compute_type = compute_type or os.environ.get("MT_COMPUTE_TYPE", "int8")
|
|
self.threads = threads or int(os.environ.get("MT_THREADS", "2"))
|
|
self._name: str | None = None
|
|
self._translator = None
|
|
self._tokenizer = None
|
|
|
|
def _load(self, name: str):
|
|
"""Keep exactly one model resident — three at once would not fit a 4GB box."""
|
|
if self._name == name:
|
|
return self._translator, self._tokenizer
|
|
import ctranslate2
|
|
import transformers
|
|
|
|
self._translator = None # free the previous model before allocating the next
|
|
path = os.path.join(self.model_root, name)
|
|
self._tokenizer = transformers.AutoTokenizer.from_pretrained(path)
|
|
self._translator = ctranslate2.Translator(
|
|
path, device="cpu", compute_type=self.compute_type, intra_threads=self.threads
|
|
)
|
|
self._name = name
|
|
return self._translator, self._tokenizer
|
|
|
|
def _hop(self, texts: list[str], model: str, token: str | None) -> list[str]:
|
|
translator, tok = self._load(model)
|
|
prepared = [f"{token} {t}" if token else t for t in texts]
|
|
batch = [tok.convert_ids_to_tokens(tok.encode(t)) for t in prepared]
|
|
results = translator.translate_batch(batch)
|
|
return [
|
|
tok.decode(tok.convert_tokens_to_ids(r.hypotheses[0]), skip_special_tokens=True)
|
|
for r in results
|
|
]
|
|
|
|
def translate(self, texts: list[str], src: str, tgt: str) -> list[str]:
|
|
if not texts:
|
|
return []
|
|
route = ROUTES.get((src, tgt))
|
|
if route is None:
|
|
raise ValueError(f"no translation route from {src} to {tgt}")
|
|
for model, token in route:
|
|
texts = self._hop(texts, model, token)
|
|
return texts
|