Compare commits
3 Commits
260ea3d383
...
7f7f535bda
| Author | SHA1 | Date | |
|---|---|---|---|
| 7f7f535bda | |||
| 5898f96edb | |||
|
|
24dbc9b961 |
@ -5,10 +5,14 @@
|
|||||||
# translation happens here, asynchronously, and every generated sibling lands
|
# translation happens here, asynchronously, and every generated sibling lands
|
||||||
# as a reviewable bot commit.
|
# as a reviewable bot commit.
|
||||||
#
|
#
|
||||||
# Translation is self-hosted — the model ships inside vienalatina/translate,
|
# Translation is self-hosted — OPUS-MT runs on CPU inside the image, so there
|
||||||
# so there is no API key and no third-party request. Build that image on the
|
# is no API key and no third-party request. Before the first run, on the server:
|
||||||
# server before the first run:
|
# bash scripts/fetch-models.sh # -> /srv/mt-models
|
||||||
# docker build -t vienalatina/translate:1 docker/translate
|
# docker build -t vienalatina/translate:2 docker/translate
|
||||||
|
#
|
||||||
|
# The models are mounted rather than baked in, so swapping them later does not
|
||||||
|
# mean rebuilding the image. The repo must be marked "trusted" in Woodpecker
|
||||||
|
# for this mount, which it already is for the deploy step.
|
||||||
#
|
#
|
||||||
# Secrets to configure in Woodpecker (repo settings → secrets):
|
# Secrets to configure in Woodpecker (repo settings → secrets):
|
||||||
# gitea_push_token — Gitea token for the translations bot user
|
# gitea_push_token — Gitea token for the translations bot user
|
||||||
@ -23,8 +27,10 @@ when:
|
|||||||
|
|
||||||
steps:
|
steps:
|
||||||
translate:
|
translate:
|
||||||
image: vienalatina/translate:1
|
image: vienalatina/translate:2
|
||||||
pull: false # built locally on the server, never fetched from a registry
|
pull: false # built locally on the server, never fetched from a registry
|
||||||
|
volumes:
|
||||||
|
- /srv/mt-models:/opt/mt/models:ro
|
||||||
environment:
|
environment:
|
||||||
GITEA_PUSH_TOKEN:
|
GITEA_PUSH_TOKEN:
|
||||||
from_secret: gitea_push_token
|
from_secret: gitea_push_token
|
||||||
|
|||||||
@ -1,8 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Über uns"
|
|
||||||
lang: de
|
|
||||||
manual_translation: true
|
|
||||||
translated_from: es
|
|
||||||
---
|
|
||||||
|
|
||||||
**Viena Latina** ist die lateinamerikanische Gemeinschaft in Wien: Tourismusführer, kulturelle Agenda, Gastronomie, Gemeinschaftsleben und lateinamerikanische Geschäfte in der Stadt.
|
|
||||||
@ -1,8 +0,0 @@
|
|||||||
---
|
|
||||||
title: Sobre
|
|
||||||
lang: pt-br
|
|
||||||
manual_translation: false
|
|
||||||
translated_from: es
|
|
||||||
---
|
|
||||||
|
|
||||||
**Viena Latina** é a comunidade latino-americana em Viena: guias de turismo, agenda cultural, gastronomia, vida comunitária e lojas latinas na cidade.
|
|
||||||
@ -1,8 +0,0 @@
|
|||||||
---
|
|
||||||
title: Kontakt
|
|
||||||
lang: de
|
|
||||||
manual_translation: false
|
|
||||||
translated_from: es
|
|
||||||
---
|
|
||||||
|
|
||||||
Schreiben Sie uns an **hola@vienalatina.com**.
|
|
||||||
@ -1,8 +0,0 @@
|
|||||||
---
|
|
||||||
title: Contacto
|
|
||||||
lang: pt-br
|
|
||||||
manual_translation: false
|
|
||||||
translated_from: es
|
|
||||||
---
|
|
||||||
|
|
||||||
Escreva-nos a **hola@vienalatina.com**.
|
|
||||||
@ -1,18 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Willkommen bei Viena Latina"
|
|
||||||
date: 2026-07-31
|
|
||||||
slug: bienvenida-a-viena-latina
|
|
||||||
lang: de
|
|
||||||
manual_translation: false
|
|
||||||
categories: [Comunidad]
|
|
||||||
description: "Die neue Website von Viena Latina: schneller, selbst gehostet und mit überprüfbaren Übersetzungen auf Deutsch und Portugiesisch."
|
|
||||||
---
|
|
||||||
|
|
||||||
Das ist das neue Zuhause von **Viena Latina**. Die Website ist jetzt komplett
|
|
||||||
statisch: Jeder Artikel wird auf Spanisch veröffentlicht und automatisch ins
|
|
||||||
Deutsche und brasilianische Portugiesisch übersetzt, mit eigenen URLs pro
|
|
||||||
Sprache.
|
|
||||||
|
|
||||||
Für unsere Leserinnen und Leser ändert sich nichts — dieselben Rubriken wie
|
|
||||||
immer: Tourismus, Kultur, Gastronomie, Gemeinschaft und Handel. Alles lädt
|
|
||||||
schneller, ohne Cookies und ohne Tracker.
|
|
||||||
@ -1,16 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Bienvenida a Viena Latina"
|
|
||||||
date: 2026-07-31
|
|
||||||
lang: es
|
|
||||||
manual_translation: false
|
|
||||||
categories: [Comunidad]
|
|
||||||
description: "El nuevo sitio de Viena Latina: más rápido, autoalojado y con traducciones revisables en alemán y portugués."
|
|
||||||
---
|
|
||||||
|
|
||||||
Este es el nuevo hogar de **Viena Latina**. El sitio ahora es completamente
|
|
||||||
estático: cada artículo se publica en español y se traduce automáticamente al
|
|
||||||
alemán y al portugués brasileño, con URLs propias para cada idioma.
|
|
||||||
|
|
||||||
Nada cambia para quienes nos leen — las mismas secciones de siempre: Turismo,
|
|
||||||
Cultura, Gastronomía, Comunidad y Comercio. Todo carga más rápido, sin cookies
|
|
||||||
y sin rastreadores.
|
|
||||||
@ -1,17 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Bem-vindos ao Viena Latina"
|
|
||||||
date: 2026-07-31
|
|
||||||
slug: bienvenida-a-viena-latina
|
|
||||||
lang: pt-br
|
|
||||||
manual_translation: false
|
|
||||||
categories: [Comunidad]
|
|
||||||
description: "O novo site do Viena Latina: mais rápido, auto-hospedado e com traduções revisáveis em alemão e português."
|
|
||||||
---
|
|
||||||
|
|
||||||
Este é o novo lar do **Viena Latina**. O site agora é totalmente estático:
|
|
||||||
cada artigo é publicado em espanhol e traduzido automaticamente para o alemão
|
|
||||||
e o português brasileiro, com URLs próprias para cada idioma.
|
|
||||||
|
|
||||||
Nada muda para quem nos lê — as mesmas seções de sempre: Turismo, Cultura,
|
|
||||||
Gastronomia, Comunidade e Comércio. Tudo carrega mais rápido, sem cookies e
|
|
||||||
sem rastreadores.
|
|
||||||
13
content/post/hola-mundo.es.md
Normal file
13
content/post/hola-mundo.es.md
Normal file
@ -0,0 +1,13 @@
|
|||||||
|
---
|
||||||
|
title: "Hola mundo"
|
||||||
|
date: 2026-09-18
|
||||||
|
lang: es
|
||||||
|
manual_translation: false
|
||||||
|
categories: [Comunidad]
|
||||||
|
---
|
||||||
|
|
||||||
|
Bienvenidos al nuevo sitio de Viena Latina.
|
||||||
|
|
||||||
|
Publicamos en español, y cada artículo aparece también en alemán y en portugués.
|
||||||
|
Aquí vas a encontrar turismo, cultura, gastronomía, comunidad y comercio latino
|
||||||
|
en Viena.
|
||||||
@ -1,21 +1,20 @@
|
|||||||
# Pipeline image for the translate step.
|
# Pipeline image for the translate step.
|
||||||
#
|
#
|
||||||
# Model and tokenizer are baked in, so a publish makes zero network calls and
|
# Models are NOT baked in — they live on the host at /srv/mt-models (see
|
||||||
# needs no API key. torch is deliberately absent: it is only required to
|
# scripts/fetch-models.sh) and are mounted read-only by .woodpecker.yml. That
|
||||||
# *convert* a model, and the default PyPI wheel drags in ~2.5GB of CUDA
|
# keeps this image small and lets a model change be a directory swap instead of
|
||||||
# libraries this CPU-only box will never use.
|
# a 1GB image rebuild, which matters because model choice turned out to need
|
||||||
|
# iteration.
|
||||||
#
|
#
|
||||||
# Build on the server (once, and again only when changing models):
|
# torch is deliberately absent: it is only needed to *convert* a model, and the
|
||||||
# docker build -t vienalatina/translate:1 docker/translate
|
# default PyPI wheel drags in ~2.5GB of CUDA libraries this CPU-only box will
|
||||||
|
# never use. Conversion happens in fetch-models.sh, not here.
|
||||||
#
|
#
|
||||||
# The default model is a pre-converted CTranslate2 build, which avoids running
|
# Build on the server:
|
||||||
# ct2-transformers-converter on a 4GB box — it gets OOM-killed there.
|
# docker build -t vienalatina/translate:2 docker/translate
|
||||||
|
|
||||||
FROM python:3.12-slim
|
FROM python:3.12-slim
|
||||||
|
|
||||||
ARG MT_MODEL_REPO=michaelfeil/ct2fast-m2m100_418M
|
|
||||||
ARG MT_TOKENIZER_REPO=facebook/m2m100_418M
|
|
||||||
|
|
||||||
# translate.py shells out to git to diff the push and commit the siblings back.
|
# translate.py shells out to git to diff the push and commit the siblings back.
|
||||||
RUN apt-get update -qq \
|
RUN apt-get update -qq \
|
||||||
&& apt-get install -qq -y --no-install-recommends git \
|
&& apt-get install -qq -y --no-install-recommends git \
|
||||||
@ -26,17 +25,9 @@ RUN pip install --no-cache-dir \
|
|||||||
transformers \
|
transformers \
|
||||||
sentencepiece \
|
sentencepiece \
|
||||||
sentencex \
|
sentencex \
|
||||||
pyyaml \
|
pyyaml
|
||||||
huggingface_hub
|
|
||||||
|
|
||||||
RUN python -c "from huggingface_hub import snapshot_download; \
|
ENV MT_PROVIDER=opus \
|
||||||
snapshot_download('${MT_MODEL_REPO}', local_dir='/opt/mt/model')" \
|
MT_MODEL_DIR=/opt/mt/models \
|
||||||
&& python -c "import transformers; \
|
|
||||||
transformers.AutoTokenizer.from_pretrained('${MT_TOKENIZER_REPO}') \
|
|
||||||
.save_pretrained('/opt/mt/tokenizer')"
|
|
||||||
|
|
||||||
# The published artifact is float16; CTranslate2 quantises to int8 on load.
|
|
||||||
ENV MT_MODEL_DIR=/opt/mt/model \
|
|
||||||
MT_TOKENIZER=/opt/mt/tokenizer \
|
|
||||||
MT_COMPUTE_TYPE=int8 \
|
MT_COMPUTE_TYPE=int8 \
|
||||||
MT_THREADS=2
|
MT_THREADS=2
|
||||||
|
|||||||
30
fetch-models.sh
Normal file
30
fetch-models.sh
Normal file
@ -0,0 +1,30 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Download and convert the OPUS-MT models into $DEST (default /srv/mt-models).
|
||||||
|
# Containers run as root so pip lands on PATH and can write to the root-owned
|
||||||
|
# target; the pipeline mounts that directory read-only.
|
||||||
|
set -euo pipefail
|
||||||
|
DEST="${1:-/srv/mt-models}"
|
||||||
|
MODELS=(
|
||||||
|
"Helsinki-NLP/opus-mt-es-de"
|
||||||
|
"Helsinki-NLP/opus-mt-tc-big-de-es"
|
||||||
|
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
|
||||||
|
)
|
||||||
|
for REPO in "${MODELS[@]}"; do
|
||||||
|
NAME="${REPO##*/}"
|
||||||
|
if [ -f "$DEST/$NAME/model.bin" ]; then
|
||||||
|
echo "== $NAME already present, skipping"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
echo "== converting $NAME"
|
||||||
|
docker run --rm --memory=3g -v "$DEST:/out" python:3.12-slim bash -c "
|
||||||
|
set -e
|
||||||
|
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
|
||||||
|
--extra-index-url https://download.pytorch.org/whl/cpu
|
||||||
|
ct2-transformers-converter --model $REPO --output_dir /out/$NAME --quantization int8
|
||||||
|
python -c 'import sys
|
||||||
|
from transformers import AutoTokenizer
|
||||||
|
AutoTokenizer.from_pretrained(sys.argv[1]).save_pretrained(sys.argv[2])' $REPO /out/$NAME
|
||||||
|
"
|
||||||
|
done
|
||||||
|
echo
|
||||||
|
du -sh "$DEST"/*
|
||||||
46
scripts/fetch-models.sh
Normal file
46
scripts/fetch-models.sh
Normal file
@ -0,0 +1,46 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Download and convert the OPUS-MT models the pipeline needs.
|
||||||
|
#
|
||||||
|
# Run once on the server; the pipeline mounts the result read-only, so changing
|
||||||
|
# models later means re-running this rather than rebuilding a 1GB image.
|
||||||
|
#
|
||||||
|
# bash scripts/fetch-models.sh # -> /srv/mt-models
|
||||||
|
# bash scripts/fetch-models.sh /tmp/mt # -> somewhere else
|
||||||
|
#
|
||||||
|
# Conversion holds the fp32 model in RAM. These are 74M-237M parameter models,
|
||||||
|
# so each peaks well under the 4GB box's headroom — unlike M2M100 418M, whose
|
||||||
|
# conversion was OOM-killed here at a 3GB cap. One container per model keeps a
|
||||||
|
# failure contained to that model.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
DEST="${1:-/srv/mt-models}"
|
||||||
|
|
||||||
|
MODELS=(
|
||||||
|
"Helsinki-NLP/opus-mt-es-de"
|
||||||
|
"Helsinki-NLP/opus-mt-tc-big-de-es"
|
||||||
|
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
|
||||||
|
)
|
||||||
|
|
||||||
|
sudo mkdir -p "$DEST"
|
||||||
|
sudo chown "$(id -u):$(id -g)" "$DEST"
|
||||||
|
|
||||||
|
for REPO in "${MODELS[@]}"; do
|
||||||
|
NAME="${REPO##*/}"
|
||||||
|
if [ -f "$DEST/$NAME/model.bin" ]; then
|
||||||
|
echo "== $NAME already converted, skipping"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
echo "== converting $NAME"
|
||||||
|
docker run --rm -u "$(id -u):$(id -g)" -e HOME=/tmp --memory=3g \
|
||||||
|
-v "$DEST:/out" python:3.12-slim bash -c "
|
||||||
|
set -e
|
||||||
|
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
|
||||||
|
--extra-index-url https://download.pytorch.org/whl/cpu
|
||||||
|
ct2-transformers-converter --model $REPO --output_dir /out/$NAME \
|
||||||
|
--quantization int8 --copy_files source.spm target.spm vocab.json tokenizer_config.json
|
||||||
|
"
|
||||||
|
done
|
||||||
|
|
||||||
|
echo
|
||||||
|
du -sh "$DEST"/*
|
||||||
|
echo "Models ready in $DEST"
|
||||||
@ -37,10 +37,18 @@ from pathlib import Path
|
|||||||
|
|
||||||
import yaml
|
import yaml
|
||||||
|
|
||||||
from translation.ctranslate_provider import CTranslate2Provider
|
|
||||||
from translation.markdown import translate_markdown, translate_text
|
from translation.markdown import translate_markdown, translate_text
|
||||||
from translation.provider import SITE_LANGS
|
from translation.provider import SITE_LANGS
|
||||||
|
|
||||||
|
|
||||||
|
def build_provider():
|
||||||
|
"""MT_PROVIDER=m2m100 falls back to the single-model engine."""
|
||||||
|
if os.environ.get("MT_PROVIDER", "opus") == "m2m100":
|
||||||
|
from translation.ctranslate_provider import CTranslate2Provider
|
||||||
|
return CTranslate2Provider()
|
||||||
|
from translation.opus_provider import OpusMTProvider
|
||||||
|
return OpusMTProvider()
|
||||||
|
|
||||||
REPO_ROOT = Path(__file__).resolve().parent.parent
|
REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
CONTENT_DIR = REPO_ROOT / "content"
|
CONTENT_DIR = REPO_ROOT / "content"
|
||||||
|
|
||||||
@ -225,7 +233,7 @@ def main() -> None:
|
|||||||
print("No authored content changed — nothing to translate.")
|
print("No authored content changed — nothing to translate.")
|
||||||
return
|
return
|
||||||
|
|
||||||
provider = CTranslate2Provider()
|
provider = build_provider()
|
||||||
written: list[Path] = []
|
written: list[Path] = []
|
||||||
for source, basename, lang in sources:
|
for source, basename, lang in sources:
|
||||||
print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):")
|
print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):")
|
||||||
|
|||||||
83
scripts/translation/opus_provider.py
Normal file
83
scripts/translation/opus_provider.py
Normal file
@ -0,0 +1,83 @@
|
|||||||
|
"""OPUS-MT (Helsinki-NLP) via CTranslate2, on CPU.
|
||||||
|
|
||||||
|
Replaces M2M100 418M, which was measured producing unusable German on real
|
||||||
|
articles — "frijoles" became "Beeren" (berries), "rompe la idea" became
|
||||||
|
"breitet die Idee" (spreads it), and the site's own name came back mangled.
|
||||||
|
|
||||||
|
Not every pair exists as a published model, so routes are explicit rather than
|
||||||
|
assumed. Verified against Hugging Face:
|
||||||
|
|
||||||
|
es -> de opus-mt-es-de (small, ~74M)
|
||||||
|
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
|
||||||
|
es <-> pt-br opus-mt-tc-big-itc-itc (Italic multilingual)
|
||||||
|
de <-> pt-br no direct model — pivots through Spanish
|
||||||
|
|
||||||
|
Marian models take the target language as a `>>xxx<<` token at the start of
|
||||||
|
the *source* text, unlike M2M100's decoder-side target_prefix.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
|
||||||
|
from .provider import Provider
|
||||||
|
|
||||||
|
# (source, target) -> [(model directory, target-language token or None), ...]
|
||||||
|
# More than one hop means a pivot translation.
|
||||||
|
ROUTES: dict[tuple[str, str], list[tuple[str, str | None]]] = {
|
||||||
|
("es", "de"): [("opus-mt-es-de", None)],
|
||||||
|
("de", "es"): [("opus-mt-tc-big-de-es", None)],
|
||||||
|
("es", "pt-br"): [("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
||||||
|
("pt-br", "es"): [("opus-mt-tc-big-itc-itc", ">>spa<<")],
|
||||||
|
("de", "pt-br"): [("opus-mt-tc-big-de-es", None),
|
||||||
|
("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
||||||
|
("pt-br", "de"): [("opus-mt-tc-big-itc-itc", ">>spa<<"),
|
||||||
|
("opus-mt-es-de", None)],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class OpusMTProvider(Provider):
|
||||||
|
def __init__(self, model_root: str | None = None, compute_type: str | None = None,
|
||||||
|
threads: int | None = None):
|
||||||
|
self.model_root = model_root or os.environ.get("MT_MODEL_DIR", "/opt/mt/models")
|
||||||
|
self.compute_type = compute_type or os.environ.get("MT_COMPUTE_TYPE", "int8")
|
||||||
|
self.threads = threads or int(os.environ.get("MT_THREADS", "2"))
|
||||||
|
self._name: str | None = None
|
||||||
|
self._translator = None
|
||||||
|
self._tokenizer = None
|
||||||
|
|
||||||
|
def _load(self, name: str):
|
||||||
|
"""Keep exactly one model resident — three at once would not fit a 4GB box."""
|
||||||
|
if self._name == name:
|
||||||
|
return self._translator, self._tokenizer
|
||||||
|
import ctranslate2
|
||||||
|
import transformers
|
||||||
|
|
||||||
|
self._translator = None # free the previous model before allocating the next
|
||||||
|
path = os.path.join(self.model_root, name)
|
||||||
|
self._tokenizer = transformers.AutoTokenizer.from_pretrained(path)
|
||||||
|
self._translator = ctranslate2.Translator(
|
||||||
|
path, device="cpu", compute_type=self.compute_type, intra_threads=self.threads
|
||||||
|
)
|
||||||
|
self._name = name
|
||||||
|
return self._translator, self._tokenizer
|
||||||
|
|
||||||
|
def _hop(self, texts: list[str], model: str, token: str | None) -> list[str]:
|
||||||
|
translator, tok = self._load(model)
|
||||||
|
prepared = [f"{token} {t}" if token else t for t in texts]
|
||||||
|
batch = [tok.convert_ids_to_tokens(tok.encode(t)) for t in prepared]
|
||||||
|
results = translator.translate_batch(batch)
|
||||||
|
return [
|
||||||
|
tok.decode(tok.convert_tokens_to_ids(r.hypotheses[0]), skip_special_tokens=True)
|
||||||
|
for r in results
|
||||||
|
]
|
||||||
|
|
||||||
|
def translate(self, texts: list[str], src: str, tgt: str) -> list[str]:
|
||||||
|
if not texts:
|
||||||
|
return []
|
||||||
|
route = ROUTES.get((src, tgt))
|
||||||
|
if route is None:
|
||||||
|
raise ValueError(f"no translation route from {src} to {tgt}")
|
||||||
|
for model, token in route:
|
||||||
|
texts = self._hop(texts, model, token)
|
||||||
|
return texts
|
||||||
Loading…
Reference in New Issue
Block a user