Compare commits
No commits in common. "7f7f535bda4ad5ac0c637558ff6fe15caa78d192" and "260ea3d3834e0fb61a10cdd2d6c0186943d0d9f9" have entirely different histories.
7f7f535bda
...
260ea3d383
@ -5,14 +5,10 @@
|
|||||||
# translation happens here, asynchronously, and every generated sibling lands
|
# translation happens here, asynchronously, and every generated sibling lands
|
||||||
# as a reviewable bot commit.
|
# as a reviewable bot commit.
|
||||||
#
|
#
|
||||||
# Translation is self-hosted — OPUS-MT runs on CPU inside the image, so there
|
# Translation is self-hosted — the model ships inside vienalatina/translate,
|
||||||
# is no API key and no third-party request. Before the first run, on the server:
|
# so there is no API key and no third-party request. Build that image on the
|
||||||
# bash scripts/fetch-models.sh # -> /srv/mt-models
|
# server before the first run:
|
||||||
# docker build -t vienalatina/translate:2 docker/translate
|
# docker build -t vienalatina/translate:1 docker/translate
|
||||||
#
|
|
||||||
# The models are mounted rather than baked in, so swapping them later does not
|
|
||||||
# mean rebuilding the image. The repo must be marked "trusted" in Woodpecker
|
|
||||||
# for this mount, which it already is for the deploy step.
|
|
||||||
#
|
#
|
||||||
# Secrets to configure in Woodpecker (repo settings → secrets):
|
# Secrets to configure in Woodpecker (repo settings → secrets):
|
||||||
# gitea_push_token — Gitea token for the translations bot user
|
# gitea_push_token — Gitea token for the translations bot user
|
||||||
@ -27,10 +23,8 @@ when:
|
|||||||
|
|
||||||
steps:
|
steps:
|
||||||
translate:
|
translate:
|
||||||
image: vienalatina/translate:2
|
image: vienalatina/translate:1
|
||||||
pull: false # built locally on the server, never fetched from a registry
|
pull: false # built locally on the server, never fetched from a registry
|
||||||
volumes:
|
|
||||||
- /srv/mt-models:/opt/mt/models:ro
|
|
||||||
environment:
|
environment:
|
||||||
GITEA_PUSH_TOKEN:
|
GITEA_PUSH_TOKEN:
|
||||||
from_secret: gitea_push_token
|
from_secret: gitea_push_token
|
||||||
|
|||||||
8
content/page/acerca.de.md
Normal file
8
content/page/acerca.de.md
Normal file
@ -0,0 +1,8 @@
|
|||||||
|
---
|
||||||
|
title: "Über uns"
|
||||||
|
lang: de
|
||||||
|
manual_translation: true
|
||||||
|
translated_from: es
|
||||||
|
---
|
||||||
|
|
||||||
|
**Viena Latina** ist die lateinamerikanische Gemeinschaft in Wien: Tourismusführer, kulturelle Agenda, Gastronomie, Gemeinschaftsleben und lateinamerikanische Geschäfte in der Stadt.
|
||||||
8
content/page/acerca.pt-br.md
Normal file
8
content/page/acerca.pt-br.md
Normal file
@ -0,0 +1,8 @@
|
|||||||
|
---
|
||||||
|
title: Sobre
|
||||||
|
lang: pt-br
|
||||||
|
manual_translation: false
|
||||||
|
translated_from: es
|
||||||
|
---
|
||||||
|
|
||||||
|
**Viena Latina** é a comunidade latino-americana em Viena: guias de turismo, agenda cultural, gastronomia, vida comunitária e lojas latinas na cidade.
|
||||||
8
content/page/contacto.de.md
Normal file
8
content/page/contacto.de.md
Normal file
@ -0,0 +1,8 @@
|
|||||||
|
---
|
||||||
|
title: Kontakt
|
||||||
|
lang: de
|
||||||
|
manual_translation: false
|
||||||
|
translated_from: es
|
||||||
|
---
|
||||||
|
|
||||||
|
Schreiben Sie uns an **hola@vienalatina.com**.
|
||||||
8
content/page/contacto.pt-br.md
Normal file
8
content/page/contacto.pt-br.md
Normal file
@ -0,0 +1,8 @@
|
|||||||
|
---
|
||||||
|
title: Contacto
|
||||||
|
lang: pt-br
|
||||||
|
manual_translation: false
|
||||||
|
translated_from: es
|
||||||
|
---
|
||||||
|
|
||||||
|
Escreva-nos a **hola@vienalatina.com**.
|
||||||
18
content/post/bienvenida-a-viena-latina.de.md
Normal file
18
content/post/bienvenida-a-viena-latina.de.md
Normal file
@ -0,0 +1,18 @@
|
|||||||
|
---
|
||||||
|
title: "Willkommen bei Viena Latina"
|
||||||
|
date: 2026-07-31
|
||||||
|
slug: bienvenida-a-viena-latina
|
||||||
|
lang: de
|
||||||
|
manual_translation: false
|
||||||
|
categories: [Comunidad]
|
||||||
|
description: "Die neue Website von Viena Latina: schneller, selbst gehostet und mit überprüfbaren Übersetzungen auf Deutsch und Portugiesisch."
|
||||||
|
---
|
||||||
|
|
||||||
|
Das ist das neue Zuhause von **Viena Latina**. Die Website ist jetzt komplett
|
||||||
|
statisch: Jeder Artikel wird auf Spanisch veröffentlicht und automatisch ins
|
||||||
|
Deutsche und brasilianische Portugiesisch übersetzt, mit eigenen URLs pro
|
||||||
|
Sprache.
|
||||||
|
|
||||||
|
Für unsere Leserinnen und Leser ändert sich nichts — dieselben Rubriken wie
|
||||||
|
immer: Tourismus, Kultur, Gastronomie, Gemeinschaft und Handel. Alles lädt
|
||||||
|
schneller, ohne Cookies und ohne Tracker.
|
||||||
16
content/post/bienvenida-a-viena-latina.es.md
Normal file
16
content/post/bienvenida-a-viena-latina.es.md
Normal file
@ -0,0 +1,16 @@
|
|||||||
|
---
|
||||||
|
title: "Bienvenida a Viena Latina"
|
||||||
|
date: 2026-07-31
|
||||||
|
lang: es
|
||||||
|
manual_translation: false
|
||||||
|
categories: [Comunidad]
|
||||||
|
description: "El nuevo sitio de Viena Latina: más rápido, autoalojado y con traducciones revisables en alemán y portugués."
|
||||||
|
---
|
||||||
|
|
||||||
|
Este es el nuevo hogar de **Viena Latina**. El sitio ahora es completamente
|
||||||
|
estático: cada artículo se publica en español y se traduce automáticamente al
|
||||||
|
alemán y al portugués brasileño, con URLs propias para cada idioma.
|
||||||
|
|
||||||
|
Nada cambia para quienes nos leen — las mismas secciones de siempre: Turismo,
|
||||||
|
Cultura, Gastronomía, Comunidad y Comercio. Todo carga más rápido, sin cookies
|
||||||
|
y sin rastreadores.
|
||||||
17
content/post/bienvenida-a-viena-latina.pt-br.md
Normal file
17
content/post/bienvenida-a-viena-latina.pt-br.md
Normal file
@ -0,0 +1,17 @@
|
|||||||
|
---
|
||||||
|
title: "Bem-vindos ao Viena Latina"
|
||||||
|
date: 2026-07-31
|
||||||
|
slug: bienvenida-a-viena-latina
|
||||||
|
lang: pt-br
|
||||||
|
manual_translation: false
|
||||||
|
categories: [Comunidad]
|
||||||
|
description: "O novo site do Viena Latina: mais rápido, auto-hospedado e com traduções revisáveis em alemão e português."
|
||||||
|
---
|
||||||
|
|
||||||
|
Este é o novo lar do **Viena Latina**. O site agora é totalmente estático:
|
||||||
|
cada artigo é publicado em espanhol e traduzido automaticamente para o alemão
|
||||||
|
e o português brasileiro, com URLs próprias para cada idioma.
|
||||||
|
|
||||||
|
Nada muda para quem nos lê — as mesmas seções de sempre: Turismo, Cultura,
|
||||||
|
Gastronomia, Comunidade e Comércio. Tudo carrega mais rápido, sem cookies e
|
||||||
|
sem rastreadores.
|
||||||
@ -1,13 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Hola mundo"
|
|
||||||
date: 2026-09-18
|
|
||||||
lang: es
|
|
||||||
manual_translation: false
|
|
||||||
categories: [Comunidad]
|
|
||||||
---
|
|
||||||
|
|
||||||
Bienvenidos al nuevo sitio de Viena Latina.
|
|
||||||
|
|
||||||
Publicamos en español, y cada artículo aparece también en alemán y en portugués.
|
|
||||||
Aquí vas a encontrar turismo, cultura, gastronomía, comunidad y comercio latino
|
|
||||||
en Viena.
|
|
||||||
@ -1,20 +1,21 @@
|
|||||||
# Pipeline image for the translate step.
|
# Pipeline image for the translate step.
|
||||||
#
|
#
|
||||||
# Models are NOT baked in — they live on the host at /srv/mt-models (see
|
# Model and tokenizer are baked in, so a publish makes zero network calls and
|
||||||
# scripts/fetch-models.sh) and are mounted read-only by .woodpecker.yml. That
|
# needs no API key. torch is deliberately absent: it is only required to
|
||||||
# keeps this image small and lets a model change be a directory swap instead of
|
# *convert* a model, and the default PyPI wheel drags in ~2.5GB of CUDA
|
||||||
# a 1GB image rebuild, which matters because model choice turned out to need
|
# libraries this CPU-only box will never use.
|
||||||
# iteration.
|
|
||||||
#
|
#
|
||||||
# torch is deliberately absent: it is only needed to *convert* a model, and the
|
# Build on the server (once, and again only when changing models):
|
||||||
# default PyPI wheel drags in ~2.5GB of CUDA libraries this CPU-only box will
|
# docker build -t vienalatina/translate:1 docker/translate
|
||||||
# never use. Conversion happens in fetch-models.sh, not here.
|
|
||||||
#
|
#
|
||||||
# Build on the server:
|
# The default model is a pre-converted CTranslate2 build, which avoids running
|
||||||
# docker build -t vienalatina/translate:2 docker/translate
|
# ct2-transformers-converter on a 4GB box — it gets OOM-killed there.
|
||||||
|
|
||||||
FROM python:3.12-slim
|
FROM python:3.12-slim
|
||||||
|
|
||||||
|
ARG MT_MODEL_REPO=michaelfeil/ct2fast-m2m100_418M
|
||||||
|
ARG MT_TOKENIZER_REPO=facebook/m2m100_418M
|
||||||
|
|
||||||
# translate.py shells out to git to diff the push and commit the siblings back.
|
# translate.py shells out to git to diff the push and commit the siblings back.
|
||||||
RUN apt-get update -qq \
|
RUN apt-get update -qq \
|
||||||
&& apt-get install -qq -y --no-install-recommends git \
|
&& apt-get install -qq -y --no-install-recommends git \
|
||||||
@ -25,9 +26,17 @@ RUN pip install --no-cache-dir \
|
|||||||
transformers \
|
transformers \
|
||||||
sentencepiece \
|
sentencepiece \
|
||||||
sentencex \
|
sentencex \
|
||||||
pyyaml
|
pyyaml \
|
||||||
|
huggingface_hub
|
||||||
|
|
||||||
ENV MT_PROVIDER=opus \
|
RUN python -c "from huggingface_hub import snapshot_download; \
|
||||||
MT_MODEL_DIR=/opt/mt/models \
|
snapshot_download('${MT_MODEL_REPO}', local_dir='/opt/mt/model')" \
|
||||||
|
&& python -c "import transformers; \
|
||||||
|
transformers.AutoTokenizer.from_pretrained('${MT_TOKENIZER_REPO}') \
|
||||||
|
.save_pretrained('/opt/mt/tokenizer')"
|
||||||
|
|
||||||
|
# The published artifact is float16; CTranslate2 quantises to int8 on load.
|
||||||
|
ENV MT_MODEL_DIR=/opt/mt/model \
|
||||||
|
MT_TOKENIZER=/opt/mt/tokenizer \
|
||||||
MT_COMPUTE_TYPE=int8 \
|
MT_COMPUTE_TYPE=int8 \
|
||||||
MT_THREADS=2
|
MT_THREADS=2
|
||||||
|
|||||||
@ -1,30 +0,0 @@
|
|||||||
#!/usr/bin/env bash
|
|
||||||
# Download and convert the OPUS-MT models into $DEST (default /srv/mt-models).
|
|
||||||
# Containers run as root so pip lands on PATH and can write to the root-owned
|
|
||||||
# target; the pipeline mounts that directory read-only.
|
|
||||||
set -euo pipefail
|
|
||||||
DEST="${1:-/srv/mt-models}"
|
|
||||||
MODELS=(
|
|
||||||
"Helsinki-NLP/opus-mt-es-de"
|
|
||||||
"Helsinki-NLP/opus-mt-tc-big-de-es"
|
|
||||||
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
|
|
||||||
)
|
|
||||||
for REPO in "${MODELS[@]}"; do
|
|
||||||
NAME="${REPO##*/}"
|
|
||||||
if [ -f "$DEST/$NAME/model.bin" ]; then
|
|
||||||
echo "== $NAME already present, skipping"
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
echo "== converting $NAME"
|
|
||||||
docker run --rm --memory=3g -v "$DEST:/out" python:3.12-slim bash -c "
|
|
||||||
set -e
|
|
||||||
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
|
|
||||||
--extra-index-url https://download.pytorch.org/whl/cpu
|
|
||||||
ct2-transformers-converter --model $REPO --output_dir /out/$NAME --quantization int8
|
|
||||||
python -c 'import sys
|
|
||||||
from transformers import AutoTokenizer
|
|
||||||
AutoTokenizer.from_pretrained(sys.argv[1]).save_pretrained(sys.argv[2])' $REPO /out/$NAME
|
|
||||||
"
|
|
||||||
done
|
|
||||||
echo
|
|
||||||
du -sh "$DEST"/*
|
|
||||||
@ -1,46 +0,0 @@
|
|||||||
#!/usr/bin/env bash
|
|
||||||
# Download and convert the OPUS-MT models the pipeline needs.
|
|
||||||
#
|
|
||||||
# Run once on the server; the pipeline mounts the result read-only, so changing
|
|
||||||
# models later means re-running this rather than rebuilding a 1GB image.
|
|
||||||
#
|
|
||||||
# bash scripts/fetch-models.sh # -> /srv/mt-models
|
|
||||||
# bash scripts/fetch-models.sh /tmp/mt # -> somewhere else
|
|
||||||
#
|
|
||||||
# Conversion holds the fp32 model in RAM. These are 74M-237M parameter models,
|
|
||||||
# so each peaks well under the 4GB box's headroom — unlike M2M100 418M, whose
|
|
||||||
# conversion was OOM-killed here at a 3GB cap. One container per model keeps a
|
|
||||||
# failure contained to that model.
|
|
||||||
|
|
||||||
set -euo pipefail
|
|
||||||
DEST="${1:-/srv/mt-models}"
|
|
||||||
|
|
||||||
MODELS=(
|
|
||||||
"Helsinki-NLP/opus-mt-es-de"
|
|
||||||
"Helsinki-NLP/opus-mt-tc-big-de-es"
|
|
||||||
"Helsinki-NLP/opus-mt-tc-big-itc-itc"
|
|
||||||
)
|
|
||||||
|
|
||||||
sudo mkdir -p "$DEST"
|
|
||||||
sudo chown "$(id -u):$(id -g)" "$DEST"
|
|
||||||
|
|
||||||
for REPO in "${MODELS[@]}"; do
|
|
||||||
NAME="${REPO##*/}"
|
|
||||||
if [ -f "$DEST/$NAME/model.bin" ]; then
|
|
||||||
echo "== $NAME already converted, skipping"
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
echo "== converting $NAME"
|
|
||||||
docker run --rm -u "$(id -u):$(id -g)" -e HOME=/tmp --memory=3g \
|
|
||||||
-v "$DEST:/out" python:3.12-slim bash -c "
|
|
||||||
set -e
|
|
||||||
pip install --quiet 'transformers[torch]' ctranslate2 sentencepiece \
|
|
||||||
--extra-index-url https://download.pytorch.org/whl/cpu
|
|
||||||
ct2-transformers-converter --model $REPO --output_dir /out/$NAME \
|
|
||||||
--quantization int8 --copy_files source.spm target.spm vocab.json tokenizer_config.json
|
|
||||||
"
|
|
||||||
done
|
|
||||||
|
|
||||||
echo
|
|
||||||
du -sh "$DEST"/*
|
|
||||||
echo "Models ready in $DEST"
|
|
||||||
@ -37,18 +37,10 @@ from pathlib import Path
|
|||||||
|
|
||||||
import yaml
|
import yaml
|
||||||
|
|
||||||
|
from translation.ctranslate_provider import CTranslate2Provider
|
||||||
from translation.markdown import translate_markdown, translate_text
|
from translation.markdown import translate_markdown, translate_text
|
||||||
from translation.provider import SITE_LANGS
|
from translation.provider import SITE_LANGS
|
||||||
|
|
||||||
|
|
||||||
def build_provider():
|
|
||||||
"""MT_PROVIDER=m2m100 falls back to the single-model engine."""
|
|
||||||
if os.environ.get("MT_PROVIDER", "opus") == "m2m100":
|
|
||||||
from translation.ctranslate_provider import CTranslate2Provider
|
|
||||||
return CTranslate2Provider()
|
|
||||||
from translation.opus_provider import OpusMTProvider
|
|
||||||
return OpusMTProvider()
|
|
||||||
|
|
||||||
REPO_ROOT = Path(__file__).resolve().parent.parent
|
REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
CONTENT_DIR = REPO_ROOT / "content"
|
CONTENT_DIR = REPO_ROOT / "content"
|
||||||
|
|
||||||
@ -233,7 +225,7 @@ def main() -> None:
|
|||||||
print("No authored content changed — nothing to translate.")
|
print("No authored content changed — nothing to translate.")
|
||||||
return
|
return
|
||||||
|
|
||||||
provider = build_provider()
|
provider = CTranslate2Provider()
|
||||||
written: list[Path] = []
|
written: list[Path] = []
|
||||||
for source, basename, lang in sources:
|
for source, basename, lang in sources:
|
||||||
print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):")
|
print(f"Translating {source.relative_to(REPO_ROOT)} (from {lang}):")
|
||||||
|
|||||||
@ -1,83 +0,0 @@
|
|||||||
"""OPUS-MT (Helsinki-NLP) via CTranslate2, on CPU.
|
|
||||||
|
|
||||||
Replaces M2M100 418M, which was measured producing unusable German on real
|
|
||||||
articles — "frijoles" became "Beeren" (berries), "rompe la idea" became
|
|
||||||
"breitet die Idee" (spreads it), and the site's own name came back mangled.
|
|
||||||
|
|
||||||
Not every pair exists as a published model, so routes are explicit rather than
|
|
||||||
assumed. Verified against Hugging Face:
|
|
||||||
|
|
||||||
es -> de opus-mt-es-de (small, ~74M)
|
|
||||||
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
|
|
||||||
es <-> pt-br opus-mt-tc-big-itc-itc (Italic multilingual)
|
|
||||||
de <-> pt-br no direct model — pivots through Spanish
|
|
||||||
|
|
||||||
Marian models take the target language as a `>>xxx<<` token at the start of
|
|
||||||
the *source* text, unlike M2M100's decoder-side target_prefix.
|
|
||||||
"""
|
|
||||||
|
|
||||||
from __future__ import annotations
|
|
||||||
|
|
||||||
import os
|
|
||||||
|
|
||||||
from .provider import Provider
|
|
||||||
|
|
||||||
# (source, target) -> [(model directory, target-language token or None), ...]
|
|
||||||
# More than one hop means a pivot translation.
|
|
||||||
ROUTES: dict[tuple[str, str], list[tuple[str, str | None]]] = {
|
|
||||||
("es", "de"): [("opus-mt-es-de", None)],
|
|
||||||
("de", "es"): [("opus-mt-tc-big-de-es", None)],
|
|
||||||
("es", "pt-br"): [("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
|
||||||
("pt-br", "es"): [("opus-mt-tc-big-itc-itc", ">>spa<<")],
|
|
||||||
("de", "pt-br"): [("opus-mt-tc-big-de-es", None),
|
|
||||||
("opus-mt-tc-big-itc-itc", ">>pob<<")],
|
|
||||||
("pt-br", "de"): [("opus-mt-tc-big-itc-itc", ">>spa<<"),
|
|
||||||
("opus-mt-es-de", None)],
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
class OpusMTProvider(Provider):
|
|
||||||
def __init__(self, model_root: str | None = None, compute_type: str | None = None,
|
|
||||||
threads: int | None = None):
|
|
||||||
self.model_root = model_root or os.environ.get("MT_MODEL_DIR", "/opt/mt/models")
|
|
||||||
self.compute_type = compute_type or os.environ.get("MT_COMPUTE_TYPE", "int8")
|
|
||||||
self.threads = threads or int(os.environ.get("MT_THREADS", "2"))
|
|
||||||
self._name: str | None = None
|
|
||||||
self._translator = None
|
|
||||||
self._tokenizer = None
|
|
||||||
|
|
||||||
def _load(self, name: str):
|
|
||||||
"""Keep exactly one model resident — three at once would not fit a 4GB box."""
|
|
||||||
if self._name == name:
|
|
||||||
return self._translator, self._tokenizer
|
|
||||||
import ctranslate2
|
|
||||||
import transformers
|
|
||||||
|
|
||||||
self._translator = None # free the previous model before allocating the next
|
|
||||||
path = os.path.join(self.model_root, name)
|
|
||||||
self._tokenizer = transformers.AutoTokenizer.from_pretrained(path)
|
|
||||||
self._translator = ctranslate2.Translator(
|
|
||||||
path, device="cpu", compute_type=self.compute_type, intra_threads=self.threads
|
|
||||||
)
|
|
||||||
self._name = name
|
|
||||||
return self._translator, self._tokenizer
|
|
||||||
|
|
||||||
def _hop(self, texts: list[str], model: str, token: str | None) -> list[str]:
|
|
||||||
translator, tok = self._load(model)
|
|
||||||
prepared = [f"{token} {t}" if token else t for t in texts]
|
|
||||||
batch = [tok.convert_ids_to_tokens(tok.encode(t)) for t in prepared]
|
|
||||||
results = translator.translate_batch(batch)
|
|
||||||
return [
|
|
||||||
tok.decode(tok.convert_tokens_to_ids(r.hypotheses[0]), skip_special_tokens=True)
|
|
||||||
for r in results
|
|
||||||
]
|
|
||||||
|
|
||||||
def translate(self, texts: list[str], src: str, tgt: str) -> list[str]:
|
|
||||||
if not texts:
|
|
||||||
return []
|
|
||||||
route = ROUTES.get((src, tgt))
|
|
||||||
if route is None:
|
|
||||||
raise ValueError(f"no translation route from {src} to {tgt}")
|
|
||||||
for model, token in route:
|
|
||||||
texts = self._hop(texts, model, token)
|
|
||||||
return texts
|
|
||||||
Loading…
Reference in New Issue
Block a user