Publishing a second "Hola mundo" through the CMS broke the first one's page:
it rendered the article, then a second copy of the entire site. Three silent
failures lined up.
Decap could not write hola-mundo.es.md twice, so it wrote hola-mundo.es-1.md.
That suffix is not a language, so Hugo stopped treating the file as a Spanish
sibling and translate.py's split_lang() skipped it — the post went live in
Spanish alone, and no German or Portuguese was ever generated. Meanwhile
permalinks used "/:slug/", and :slug falls back to the title, so both files
claimed /hola-mundo/; Hugo wrote both documents into that one index.html.
Nothing failed. The pipeline was green throughout.
Each layer now refuses its part: post permalinks carry the year and month, the
CMS prefixes new filenames with the date, and translate.py aborts on a name
ending in a clash counter rather than quietly declining to translate it. The
build also runs with --printPathWarnings --panicOnWarning, so any future pair
of pages targeting one path fails the build instead of corrupting the output.
Existing post URLs change shape (/hola-mundo/ becomes /2026/09/hola-mundo/).
That costs nothing today, with one real post and no inbound links, and gets
expensive to change later.
Verified: 16/16 pipeline checks and 15/15 markdown checks still pass, the
clash-counter name aborts with the rename instruction, and an ordinary
filename still parses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
M2M100 418M failed the quality gate on real articles. Measured on a 420-word
post: "frijoles" came back as "Beeren" (berries), "rompe la idea" as "breitet
die Idee" (spreads it — the opposite), "así sabe mi barrio" as "so weiß mein
Viertel" (knows, not tastes), and the opening sentence was not grammatical
German. Invented non-words throughout ("verforscht", "Treffenraum").
The pipeline itself was never at fault — masking held, structure survived,
zero placeholder failures across 14 files. The model was.
OPUS-MT is stronger on these specific pairs, and fits the existing CX22 with
no server upgrade, which was the constraint. Licence moves from MIT to
CC-BY-4.0, so attribution now ships with the platform.
Routes are verified against Hugging Face rather than assumed — the earlier
research could not reach HF, and half the names it guessed do not exist:
es -> de opus-mt-es-de (small, ~74M)
de -> es opus-mt-tc-big-de-es (tc-big, ~237M)
es <-> pt-br opus-mt-tc-big-itc-itc (>>pob<< / >>spa<<)
de <-> pt-br no model in either direction — pivots through Spanish
Models now live at /srv/mt-models and are mounted read-only rather than baked
into the image. Model choice has needed iteration, and a directory swap beats
a 1GB image rebuild each time. It also keeps model conversion out of the image
build, which matters on a box that OOM-killed the last conversion.
MT_PROVIDER=m2m100 still selects the old engine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
Measured against M2M100 418M on real content: the OpenNMT ⦅0⦆ convention was
dropped on every single occurrence of a protected term. SentencePiece
fragments punctuation runs, and the model then has nothing it recognises as a
unit to copy across.
The brand paid for it. "Viena Latina" came back as "Wien Latin" in one file
and "Vienna Latina" in another — the same name rendered two different wrong
ways across two pages, which is worse than being consistently wrong. "Grätzl"
survived untouched throughout, because it is genuinely unknown to the model,
whereas "Viena" reads as a city name and gets translated.
Masks are now word-shaped (Zq0Xv), which a model treats like an unknown proper
noun and carries through rather than translating. Matching is case- and
space-insensitive, since models re-case and pad these.
The tiered fallback stays as the safety net for when it is still dropped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
First real pipeline run died with:
PlaceholderError: masked span 'Viena Latina' came back 0 times
M2M100 drops placeholder tokens often enough that failing the pipeline on
mismatch would block the whole site deploy over a single proper noun. The
verification itself was right — it stopped a literal ⦅0⦆ reaching a
reader — but the policy was too blunt.
Masks are now tiered by how much they actually matter. Markup must survive;
terminology is a preference. So: try markup + terms, and on a lost term retry
guarding only markup, accepting the term may come back translated. Only if
markup itself is lost does the segment stay in the source language.
A stray placeholder or mangled URL still never reaches a reader, but one
awkward proper noun no longer blocks a publish.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn
The DeepL free tier is metered (it failed in production with HTTP 456
Quota exceeded) and would require every deployment of this platform to
carry its own API account. Translation now runs on M2M100 418M (MIT) via
CTranslate2, shipped inside the pipeline image: no key, no quota, and no
content or visitor data leaving the server.
Also restores the multi-source behaviour of the original WordPress plugin,
which the Python port had narrowed to Spanish-only. Any of the three site
languages can now be the authored original.
Because any language can be a source, loop prevention is no longer
structural and is now explicit: generated siblings carry `translated_from`
and are never treated as sources, and the bot's own [skip-translate]
commits are skipped outright (that marker was already being written but
never read).
Markup protection moves in-process now that DeepL's tag_handling=html is
gone. Code blocks and raw HTML pass through untouched; link targets,
inline code and protected community terms are masked with placeholders
that are verified to survive the round trip, failing the pipeline rather
than shipping corrupted text.
Two fixes along the way:
- Generated siblings no longer inherit the source's `slug`. They did,
which meant the first retranslation of a WordPress-migrated post moved
/de/<german-slug>/ onto /de/<spanish-slug>/ and destroyed the inbound
link preservation wp-to-hugo.py exists for.
- `manual_translation` now works from the CMS. Decap only ever exposed it
on the source while the script read it on the target, so the toggle did
nothing. It now means "hands off" on both sides.
wp-to-hugo.py marks migrated Polylang siblings frozen, since those are
human translations and regenerating them would replace them with weaker
machine output.
Adds --backfill for sources missing siblings, which also fixes the
existing 404s on /de/page/acerca/ and /pt-br/page/contacto/.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NizVpJ2dwzCbjCrTLCjeHn