The DeepL free tier is metered (it failed in production with HTTP 456 Quota exceeded) and would require every deployment of this platform to carry its own API account. Translation now runs on M2M100 418M (MIT) via CTranslate2, shipped inside the pipeline image: no key, no quota, and no content or visitor data leaving the server. Also restores the multi-source behaviour of the original WordPress plugin, which the Python port had narrowed to Spanish-only. Any of the three site languages can now be the authored original. Because any language can be a source, loop prevention is no longer structural and is now explicit: generated siblings carry `translated_from` and are never treated as sources, and the bot's own [skip-translate] commits are skipped outright (that marker was already being written but never read). Markup protection moves in-process now that DeepL's tag_handling=html is gone. Code blocks and raw HTML pass through untouched; link targets, inline code and protected community terms are masked with placeholders that are verified to survive the round trip, failing the pipeline rather than shipping corrupted text. Two fixes along the way: - Generated siblings no longer inherit the source's `slug`. They did, which meant the first retranslation of a WordPress-migrated post moved /de/<german-slug>/ onto /de/<spanish-slug>/ and destroyed the inbound link preservation wp-to-hugo.py exists for. - `manual_translation` now works from the CMS. Decap only ever exposed it on the source while the script read it on the target, so the toggle did nothing. It now means "hands off" on both sides. wp-to-hugo.py marks migrated Polylang siblings frozen, since those are human translations and regenerating them would replace them with weaker machine output. Adds --backfill for sources missing siblings, which also fixes the existing 404s on /de/page/acerca/ and /pt-br/page/contacto/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
43 lines
1.5 KiB
Docker
43 lines
1.5 KiB
Docker
# Pipeline image for the translate step.
|
|
#
|
|
# Model and tokenizer are baked in, so a publish makes zero network calls and
|
|
# needs no API key. torch is deliberately absent: it is only required to
|
|
# *convert* a model, and the default PyPI wheel drags in ~2.5GB of CUDA
|
|
# libraries this CPU-only box will never use.
|
|
#
|
|
# Build on the server (once, and again only when changing models):
|
|
# docker build -t vienalatina/translate:1 docker/translate
|
|
#
|
|
# The default model is a pre-converted CTranslate2 build, which avoids running
|
|
# ct2-transformers-converter on a 4GB box — it gets OOM-killed there.
|
|
|
|
FROM python:3.12-slim
|
|
|
|
ARG MT_MODEL_REPO=michaelfeil/ct2fast-m2m100_418M
|
|
ARG MT_TOKENIZER_REPO=facebook/m2m100_418M
|
|
|
|
# translate.py shells out to git to diff the push and commit the siblings back.
|
|
RUN apt-get update -qq \
|
|
&& apt-get install -qq -y --no-install-recommends git \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
RUN pip install --no-cache-dir \
|
|
ctranslate2 \
|
|
transformers \
|
|
sentencepiece \
|
|
sentencex \
|
|
pyyaml \
|
|
huggingface_hub
|
|
|
|
RUN python -c "from huggingface_hub import snapshot_download; \
|
|
snapshot_download('${MT_MODEL_REPO}', local_dir='/opt/mt/model')" \
|
|
&& python -c "import transformers; \
|
|
transformers.AutoTokenizer.from_pretrained('${MT_TOKENIZER_REPO}') \
|
|
.save_pretrained('/opt/mt/tokenizer')"
|
|
|
|
# The published artifact is float16; CTranslate2 quantises to int8 on load.
|
|
ENV MT_MODEL_DIR=/opt/mt/model \
|
|
MT_TOKENIZER=/opt/mt/tokenizer \
|
|
MT_COMPUTE_TYPE=int8 \
|
|
MT_THREADS=2
|