Compare commits

...

3 Commits

Author SHA1 Message Date
7f7f535bda Reset content; start from a single post
All checks were successful
ci/woodpecker/push/woodpecker Pipeline was successful
2026-09-18 13:25:23 +00:00
5898f96edb Merge branch 'claude/relaxed-faraday-h4zd09' of https://github.com/pablovolenski/vienalatina 2026-09-18 12:43:26 +00:00
Claude
24dbc9b961
Switch translation to OPUS-MT; mount models instead of baking them
M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").

The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.

OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.

Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:

  es -> de       opus-mt-es-de           (small, ~74M)
  de -> es       opus-mt-tc-big-de-es    (tc-big, ~237M)
  es <-> pt-br   opus-mt-tc-big-itc-itc  (>>pob<< / >>spa<<)
  de <-> pt-br   no model in either direction — pivots through Spanish

Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.

MT_PROVIDER=m2m100 still selects the old engine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
2026-09-18 12:41:45 +00:00
14 changed files with 206 additions and 112 deletions

View File

@ -5,10 +5,14 @@
# translation happens here, asynchronously, and every generated sibling lands # translation happens here, asynchronously, and every generated sibling lands
# as a reviewable bot commit. # as a reviewable bot commit.
# #
# Translation is self-hosted — the model ships inside vienalatina/translate, # Translation is self-hosted — OPUS-MT runs on CPU inside the image, so there
# so there is no API key and no third-party request. Build that image on the # is no API key and no third-party request. Before the first run, on the server:
# server before the first run: # bash scripts/fetch-models.sh # -> /srv/mt-models
# docker build -t vienalatina/translate:1 docker/translate # docker build -t vienalatina/translate:2 docker/translate
#
# The models are mounted rather than baked in, so swapping them later does not
# mean rebuilding the image. The repo must be marked "trusted" in Woodpecker
# for this mount, which it already is for the deploy step.
# #
# Secrets to configure in Woodpecker (repo settings → secrets): # Secrets to configure in Woodpecker (repo settings → secrets):
# gitea_push_token — Gitea token for the translations bot user # gitea_push_token — Gitea token for the translations bot user
@ -23,8 +27,10 @@ when:
steps: steps:
translate: translate:
image: vienalatina/translate:1 image: vienalatina/translate:2
pull: false # built locally on the server, never fetched from a registry pull: false # built locally on the server, never fetched from a registry
volumes:
- /srv/mt-models:/opt/mt/models:ro
environment: environment:
GITEA_PUSH_TOKEN: GITEA_PUSH_TOKEN:
from_secret: gitea_push_token from_secret: gitea_push_token

View File

@ -1,8 +0,0 @@
---
title: "Über uns"
lang: de
manual_translation: true
translated_from: es
---
**Viena Latina** ist die lateinamerikanische Gemeinschaft in Wien: Tourismusführer, kulturelle Agenda, Gastronomie, Gemeinschaftsleben und lateinamerikanische Geschäfte in der Stadt.

View File

@ -1,8 +0,0 @@
---
title: Sobre
lang: pt-br
manual_translation: false
translated_from: es
---
**Viena Latina** é a comunidade latino-americana em Viena: guias de turismo, agenda cultural, gastronomia, vida comunitária e lojas latinas na cidade.

View File

@ -1,8 +0,0 @@
---
title: Kontakt
lang: de
manual_translation: false
translated_from: es
---
Schreiben Sie uns an **hola@vienalatina.com**.

View File

@ -1,8 +0,0 @@
---
title: Contacto
lang: pt-br
manual_translation: false
translated_from: es
---
Escreva-nos a **hola@vienalatina.com**.

View File

@ -1,18 +0,0 @@
---
title: "Willkommen bei Viena Latina"
date: 2026-07-31
slug: bienvenida-a-viena-latina
lang: de
manual_translation: false
categories: [Comunidad]
description: "Die neue Website von Viena Latina: schneller, selbst gehostet und mit überprüfbaren Übersetzungen auf Deutsch und Portugiesisch."
---
Das ist das neue Zuhause von **Viena Latina**. Die Website ist jetzt komplett
statisch: Jeder Artikel wird auf Spanisch veröffentlicht und automatisch ins
Deutsche und brasilianische Portugiesisch übersetzt, mit eigenen URLs pro
Sprache.
Für unsere Leserinnen und Leser ändert sich nichts — dieselben Rubriken wie
immer: Tourismus, Kultur, Gastronomie, Gemeinschaft und Handel. Alles lädt
schneller, ohne Cookies und ohne Tracker.

View File

@ -1,16 +0,0 @@
---
title: "Bienvenida a Viena Latina"
date: 2026-07-31
lang: es
manual_translation: false
categories: [Comunidad]
description: "El nuevo sitio de Viena Latina: más rápido, autoalojado y con traducciones revisables en alemán y portugués."
---
Este es el nuevo hogar de **Viena Latina**. El sitio ahora es completamente
estático: cada artículo se publica en español y se traduce automáticamente al
alemán y al portugués brasileño, con URLs propias para cada idioma.
Nada cambia para quienes nos leen — las mismas secciones de siempre: Turismo,
Cultura, Gastronomía, Comunidad y Comercio. Todo carga más rápido, sin cookies
y sin rastreadores.

View File

@ -1,17 +0,0 @@
---
title: "Bem-vindos ao Viena Latina"
date: 2026-07-31
slug: bienvenida-a-viena-latina
lang: pt-br
manual_translation: false
categories: [Comunidad]
description: "O novo site do Viena Latina: mais rápido, auto-hospedado e com traduções revisáveis em alemão e português."
---
Este é o novo lar do **Viena Latina**. O site agora é totalmente estático:
cada artigo é publicado em espanhol e traduzido automaticamente para o alemão
e o português brasileiro, com URLs próprias para cada idioma.
Nada muda para quem nos lê — as mesmas seções de sempre: Turismo, Cultura,
Gastronomia, Comunidade e Comércio. Tudo carrega mais rápido, sem cookies e
sem rastreadores.

View File

@ -0,0 +1,13 @@
---
title: "Hola mundo"
date: 2026-09-18
lang: es
manual_translation: false
categories: [Comunidad]
---
Bienvenidos al nuevo sitio de Viena Latina.
Publicamos en español, y cada artículo aparece también en alemán y en portugués.
Aquí vas a encontrar turismo, cultura, gastronomía, comunidad y comercio latino
en Viena.

View File

@ -1,21 +1,20 @@
# Pipeline image for the translate step. # Pipeline image for the translate step.
# #
# Model and tokenizer are baked in, so a publish makes zero network calls and # Models are NOT baked in — they live on the host at /srv/mt-models (see
# needs no API key. torch is deliberately absent: it is only required to # scripts/fetch-models.sh) and are mounted read-only by .woodpecker.yml. That
# *convert* a model, and the default PyPI wheel drags in ~2.5GB of CUDA # keeps this image small and lets a model change be a directory swap instead of
# libraries this CPU-only box will never use. # a 1GB image rebuild, which matters because model choice turned out to need
# iteration.
# #
# Build on the server (once, and again only when changing models): # torch is deliberately absent: it is only needed to *convert* a model, and the
# docker build -t vienalatina/translate:1 docker/translate # default PyPI wheel drags in ~2.5GB of CUDA libraries this CPU-only box will
# never use. Conversion happens in fetch-models.sh, not here.
# #
# The default model is a pre-converted CTranslate2 build, which avoids running # Build on the server:
# ct2-transformers-converter on a 4GB box — it gets OOM-killed there. # docker build -t vienalatina/translate:2 docker/translate
FROM python:3.12-slim FROM python:3.12-slim
ARG MT_MODEL_REPO=michaelfeil/ct2fast-m2m100_418M
ARG MT_TOKENIZER_REPO=facebook/m2m100_418M
# translate.py shells out to git to diff the push and commit the siblings back. # translate.py shells out to git to diff the push and commit the siblings back.
RUN apt-get update -qq \ RUN apt-get update -qq \
&& apt-get install -qq -y --no-install-recommends git \ && apt-get install -qq -y --no-install-recommends git \
@ -26,17 +25,9 @@ RUN pip install --no-cache-dir \
transformers \ transformers \
sentencepiece \ sentencepiece \
sentencex \ sentencex \
pyyaml \ pyyaml
huggingface_hub
RUN python -c "from huggingface_hub import snapshot_download; \ ENV MT_PROVIDER=opus \
snapshot_download('${MT_MODEL_REPO}', local_dir='/opt/mt/model')" \ MT_MODEL_DIR=/opt/mt/models \
&& python -c "import transformers; \
transformers.AutoTokenizer.from_pretrained('${MT_TOKENIZER_REPO}') \
.save_pretrained('/opt/mt/tokenizer')"
# The published artifact is float16; CTranslate2 quantises to int8 on load.
ENV MT_MODEL_DIR=/opt/mt/model \
MT_TOKENIZER=/opt/mt/tokenizer \
MT_COMPUTE_TYPE=int8 \ MT_COMPUTE_TYPE=int8 \
MT_THREADS=2 MT_THREADS=2

30
fetch-models.sh Normal file
View File

@ -0,0 +1,30 @@
#!/usr/bin/env bash
# Download and convert the OPUS-MT models into $DEST (default /srv/mt-models).
# Containers run as root so pip lands on PATH and can write to the root-owned
# target; the pipeline mounts that directory read-only.
set -euo pipefail
DEST="${1:-/srv/mt-models}"
MODELS=(
"Helsinki-NLP/opus-mt-es-de"
"Helsinki-NLP/opus-mt-tc-big-de-es"
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
)
for REPO in "${MODELS[@]}"; do
NAME="${REPO##*/}"
if [ -f "$DEST/$NAME/model.bin" ]; then
echo "== $NAME already present, skipping"
continue
fi
echo "== converting $NAME"
docker run --rm --memory=3g -v "$DEST:/out" python:3.12-slim bash -c "
set -e
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
--extra-index-url https://download.pytorch.org/whl/cpu
ct2-transformers-converter --model $REPO --output_dir /out/$NAME --quantization int8
python -c 'import sys
from transformers import AutoTokenizer
AutoTokenizer.from_pretrained(sys.argv[1]).save_pretrained(sys.argv[2])' $REPO /out/$NAME
"
done
echo
du -sh "$DEST"/*

46
scripts/fetch-models.sh Normal file
View File

@ -0,0 +1,46 @@
#!/usr/bin/env bash
# Download and convert the OPUS-MT models the pipeline needs.
#
# Run once on the server; the pipeline mounts the result read-only, so changing
# models later means re-running this rather than rebuilding a 1GB image.
#
# bash scripts/fetch-models.sh # -> /srv/mt-models
# bash scripts/fetch-models.sh /tmp/mt # -> somewhere else
#
# Conversion holds the fp32 model in RAM. These are 74M-237M parameter models,
# so each peaks well under the 4GB box's headroom — unlike M2M100 418M, whose
# conversion was OOM-killed here at a 3GB cap. One container per model keeps a
# failure contained to that model.
set -euo pipefail
DEST="${1:-/srv/mt-models}"
MODELS=(
"Helsinki-NLP/opus-mt-es-de"
"Helsinki-NLP/opus-mt-tc-big-de-es"
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
)
sudo mkdir -p "$DEST"
sudo chown "$(id -u):$(id -g)" "$DEST"
for REPO in "${MODELS[@]}"; do
NAME="${REPO##*/}"
if [ -f "$DEST/$NAME/model.bin" ]; then
echo "== $NAME already converted, skipping"
continue
fi
echo "== converting $NAME"
docker run --rm -u "$(id -u):$(id -g)" -e HOME=/tmp --memory=3g \
-v "$DEST:/out" python:3.12-slim bash -c "
set -e
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
--extra-index-url https://download.pytorch.org/whl/cpu
ct2-transformers-converter --model $REPO --output_dir /out/$NAME \
--quantization int8 --copy_files source.spm target.spm vocab.json tokenizer_config.json
"
done
echo
du -sh "$DEST"/*
echo "Models ready in $DEST"

View File

@ -37,10 +37,18 @@ from pathlib import Path
import yaml import yaml
from translation.ctranslate_provider import CTranslate2Provider
from translation.markdown import translate_markdown, translate_text from translation.markdown import translate_markdown, translate_text
from translation.provider import SITE_LANGS from translation.provider import SITE_LANGS
def build_provider():
"""MT_PROVIDER=m2m100 falls back to the single-model engine."""
if os.environ.get("MT_PROVIDER", "opus") == "m2m100":
from translation.ctranslate_provider import CTranslate2Provider
return CTranslate2Provider()
from translation.opus_provider import OpusMTProvider
return OpusMTProvider()
REPO_ROOT = Path(__file__).resolve().parent.parent REPO_ROOT = Path(__file__).resolve().parent.parent
CONTENT_DIR = REPO_ROOT / "content" CONTENT_DIR = REPO_ROOT / "content"
@ -225,7 +233,7 @@ def main() -> None:
print("No authored content changed — nothing to translate.") print("No authored content changed — nothing to translate.")
return return
provider = CTranslate2Provider() provider = build_provider()
written: list[Path] = [] written: list[Path] = []
for source, basename, lang in sources: for source, basename, lang in sources:
print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):") print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):")

View File

@ -0,0 +1,83 @@
"""OPUS-MT (Helsinki-NLP) via CTranslate2, on CPU.
Replaces M2M100 418M, which was measured producing unusable German on real
articles — "frijoles" became "Beeren" (berries), "rompe la idea" became
"breitet die Idee" (spreads it), and the site's own name came back mangled.
Not every pair exists as a published model, so routes are explicit rather than
assumed. Verified against Hugging Face:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (Italic multilingual)
de <-> pt-br no direct model — pivots through Spanish
Marian models take the target language as a `>>xxx<<` token at the start of
the *source* text, unlike M2M100's decoder-side target_prefix.
"""
from __future__ import annotations
import os
from .provider import Provider
# (source, target) -> [(model directory, target-language token or None), ...]
# More than one hop means a pivot translation.
ROUTES: dict[tuple[str, str], list[tuple[str, str | None]]] = {
("es", "de"): [("opus-mt-es-de", None)],
("de", "es"): [("opus-mt-tc-big-de-es", None)],
("es", "pt-br"): [("opus-mt-tc-big-itc-itc", ">>pob<<")],
("pt-br", "es"): [("opus-mt-tc-big-itc-itc", ">>spa<<")],
("de", "pt-br"): [("opus-mt-tc-big-de-es", None),
("opus-mt-tc-big-itc-itc", ">>pob<<")],
("pt-br", "de"): [("opus-mt-tc-big-itc-itc", ">>spa<<"),
("opus-mt-es-de", None)],
}
class OpusMTProvider(Provider):
def __init__(self, model_root: str | None = None, compute_type: str | None = None,
threads: int | None = None):
self.model_root = model_root or os.environ.get("MT_MODEL_DIR", "/opt/mt/models")
self.compute_type = compute_type or os.environ.get("MT_COMPUTE_TYPE", "int8")
self.threads = threads or int(os.environ.get("MT_THREADS", "2"))
self._name: str | None = None
self._translator = None
self._tokenizer = None
def _load(self, name: str):
"""Keep exactly one model resident — three at once would not fit a 4GB box."""
if self._name == name:
return self._translator, self._tokenizer
import ctranslate2
import transformers
self._translator = None # free the previous model before allocating the next
path = os.path.join(self.model_root, name)
self._tokenizer = transformers.AutoTokenizer.from_pretrained(path)
self._translator = ctranslate2.Translator(
path, device="cpu", compute_type=self.compute_type, intra_threads=self.threads
)
self._name = name
return self._translator, self._tokenizer
def _hop(self, texts: list[str], model: str, token: str | None) -> list[str]:
translator, tok = self._load(model)
prepared = [f"{token} {t}" if token else t for t in texts]
batch = [tok.convert_ids_to_tokens(tok.encode(t)) for t in prepared]
results = translator.translate_batch(batch)
return [
tok.decode(tok.convert_tokens_to_ids(r.hypotheses[0]), skip_special_tokens=True)
for r in results
]
def translate(self, texts: list[str], src: str, tgt: str) -> list[str]:
if not texts:
return []
route = ROUTES.get((src, tgt))
if route is None:
raise ValueError(f"no translation route from {src} to {tgt}")
for model, token in route:
texts = self._hop(texts, model, token)
return texts