Eval: the language pack (translation with a number lock, speech, ASR)
All 24 EU official languages (27 Sep 2026): section 10 has the model per language, its FLORES-200 scores, the quality bar, and the number lock in the added languages.
Run 27 Sep 2026 on our server, branch the pre-release branch. Models on GPU0: Hy-MT2-7B (vLLM 0.29.0, decosa-lang-mt),
Qwen3-ASR-1.7B and VoxCPM2 (decosa-lang-speech). Runners: scripts/lang_eval_flores.py, lang_eval_comet.py,
lang_eval_drift.py, lang_eval_refs.py, lang_eval_laysummary.py, lang_eval_speech.py. Results in
docs/evals/language-pack/. Numbers were measured with other evals and a sibling block sharing GPU0, so latencies are
under load.
Models and licences (checked 27 Sep 2026)
| Model | Revision | Licence | Weights root (sha256 over the weight files) |
|---|---|---|---|
| tencent/Hy-MT2-7B | 9b0eb4e8f001def3e5ff6469a0ac96fdb39ec223 | Apache-2.0 (LICENSE.txt in the repo; no territory clause, unlike Hunyuan-MT) | see decosa_api/lang/data/models.json |
| Qwen/Qwen3-ASR-1.7B | 7278e1e70fe206f11671096ffdd38061171dd6e5 | Apache-2.0 (model card) | idem |
| openbmb/VoxCPM2 | 32279effe8c19989596f05d353d1447f51d9e915 | Apache-2.0 (model card; GitHub repo OpenBMB/VoxCPM Apache-2.0; no LICENSE file in the weights repo) | idem |
| Qwen/Qwen3.8-27B (the fifteen added EU languages, via the gateway; section 10) | 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 | Apache-2.0 (model card) | the gateway's receipt carries it |
| Unbabel/wmt22-comet-da (eval only) | as cached | Apache-2.0 | - |
The local weight files were hashed and match the Hugging Face LFS sha256 at those revisions. Qwen3-TTS was not used: its 10 languages lack Dutch and Polish.
Datasets: FLORES-200 (CC-BY-SA 4.0; used for scoring, not redistributed: per-sentence outputs stay outside the repo),
FLEURS (CC-BY-4.0, google/fleurs @ 70bb2e84, first 100 test sentences per language, first recording of each), and our own
synthetic sentences (drift-sources.json, CC0).
GPU memory and latency
| Part | Measured |
|---|---|
Hy-MT2-7B, vLLM --gpu-memory-utilization 0.18 --max-model-len 8192 --enforce-eager (weights 14.0 GiB, KV 1.98 GiB = 16k tokens) |
18.3 GB |
| Qwen3-ASR-1.7B + VoxCPM2 in one process, both loaded (each unloads after 15 min idle) | 10.4-11.5 GB |
| All three | 29.7 GB (ceiling 30 GB) |
vLLM 0.29 serves HunYuanDenseV1 through its Transformers backend; CUDA-graph capture fails there, so the unit runs eager. Single-stream MT: about 2 s for a 75-token sentence; FLORES at 12 in flight: p50 1.5-4.0 s per sentence (rising through the run as other evals loaded the card). ASR real-time factor 0.12-0.29; TTS 0.32-0.43 (both without batching).
1. Translation quality: FLORES-200 devtest, English into six languages (1,012 sentences each)
Block prompt (Hy-MT2's own), greedy, repetition penalty 1.05. sacrebleu 2.6.0 (chrF2++ nrefs:1|case:mixed|eff:yes|nc:6|nw:2).
| chrF++ | BLEU | COMET-22 | p50 / p95 ms | |
|---|---|---|---|---|
| German | 62.98 | 38.23 | 89.02 | 1,516 / 2,527 |
| French | 67.47 | 46.56 | 89.04 | 1,500 / 2,476 |
| Spanish | 55.19 | 29.64 | 87.59 | 1,825 / 3,235 |
| Italian | 56.53 | 29.56 | 89.64 | 2,400 / 4,224 |
| Dutch | 55.42 | 27.33 | 88.63 | 2,969 / 4,760 |
| Polish | 50.10 | 23.05 | 90.54 | 4,021 / 5,974 |
No empty outputs. Same harness with Qwen3.8-27B on 23 Sep (internal wiki page 22): German 62.80 / 88.61, French 68.77 / 88.89, Spanish 55.32 / 87.24 (chrF++ / COMET). Hy-MT2-7B is at the 27B model's level at a quarter of the memory; Italian, Dutch and Polish were not run on Qwen3.8, so no side-by-side there.
2. Planted number drift (the Decosa angle)
90 synthetic source sentences written for this eval (labels, dosing, lay-summary lines, specifications), split 30 dev / 60 test before any checker work. Hy-MT2 translated each into six languages. For each clean first-pass translation, one variant per applicable error type was made by regexes in the runner (not the checker's tables). Caught = a new high- or medium-severity flag the clean translation did not have.
Test split, checker frozen (run once):
| Error planted | Caught |
|---|---|
| a digit changed | 318 / 318 |
| a number x10 | 292 / 292 |
| a thousands group written the English way in a comma-decimal language (1.500 -> 1,500) | 10 / 10 |
| a unit swapped (mg/g, ml/l, °C/°F, cm/mm, V/W, hours/days/weeks/months/years) | 218 / 226 |
| a number and its unit dropped | 314 / 318 |
| the negation removed | 114 / 116 |
| a date changed | 36 / 36 |
| a code changed (model, EAN, phone) | 12 / 12 |
| all | 1,314 / 1,328 (98.9%) |
By language: de 217/219, fr 220/223, es 220/221, it 219/222, nl 222/223, pl 216/220.
False flags on the clean translations (test): 9 of 360 sentences flagged. Judged by me: 2 real (a French warning translated as "Contraire aux enfants...", garbled; a Spanish "2,000 W" left in English format, which reads as 2 W under Spanish conventions) and 7 false (the unit words "inches"/"Zoll"/"pollici"... missing from the tables x5, "Betriebsstunden", and "Vermeiden Sie die Mischung" for "do not mix", a negation paraphrased away). False-flag rate 7/360 (1.9%) of sentences.
After the test run (fixes informed by the test misses and by the lay-summary runs below, so no longer held out): 1,320 / 1,328 caught, 3 clean sentences flagged. The 8 remaining misses are all one case: "Children aged 8 to 14" has no unit in English, the translator added "Jahren", the plant changed it to "Monaten", and there is nothing in the source to compare with.
Dev split (used while building): 602/622 before the dev fixes, 620/621 after.
3. The same check on professional translations (FLORES-200)
A flag on a human translation is either a false alarm or a liberty the translator took. Sentences flagged of 1,012 (250 have a number, date or code in English), frozen checker -> final checker:
| de | fr | es | it | nl | pl |
|---|---|---|---|---|---|
| 49 -> 46 | 48 -> 47 | 63 -> 60 | 56 -> 55 | 68 -> 65 | 66 -> 58 |
Tuned on the FLORES dev split (997 sentences, 7-14% flagged before tuning), reported on devtest. Reading a sample of the dev flags: most are real liberties in Wikipedia-style text (miles converted to km, "300,948 sq mi" dropped, "2,220" written "2.200", "twenty" written as a compound word the tables do not read, "not" paraphrased with a prefix such as "ungefährlich"); the rest are checker limits listed at the end. FLORES is literary; label and dosing text is the target.
4. Lay summary Member-State versions (real trial results)
The final English drafts of the lay summary eval (use case 66: 12 test trials and 5 test2 trials, real ClinicalTrials.gov
results) translated by /laysummary/translate into de, fr, es, it, nl, pl. Each translated number is checked against the
English sentence and traced to the results cells in that language.
Test (12 trials, 775 sentences per language), frozen code:
| numbers traced | wrong | sentences fail / check | |
|---|---|---|---|
| de | 658 / 684 | 6 | 7 / 27 |
| fr | 665 / 704 | 13 | 13 / 35 |
| es | 695 / 750 | 15 | 17 / 39 |
| it | 659 / 696 | 12 | 23 / 27 |
| nl | 661 / 679 | 8 | 21 / 16 |
| pl | 664 / 692 | 8 | 9 / 26 |
| all | 4,002 / 4,205 (95.2%) | 62 |
Reading every failing sentence: almost all were the tracer misreading other languages ("249 des 325", "105 de las 110"
as out-of phrases; a dose inside a group name; "moins de 1 %" as "less than"; clause words "et", "y", "en"; "28 derniers
jours" as a duration; "whether or not" as a negation). Fixed, then the stored translations re-checked with the new code
(laysummary-recheck-*.json, no longer held out): 4,014 / 4,077 traced; failing sentences de 0, fr 0, es 2, it 0, nl 2,
pl 4. Test2 (5 trials): de 0, fr 0, es 0, it 0, nl 1, pl 7.
Real translation errors found: 6 Polish sentences in test2 where Hy-MT2 dropped "0 out of" ("In Part 1, 0 out of 89 people died" became "W części 1 z 89 osób zmarło", which reads "in part 1, of 89 people died"), and one Spanish "30,000 por microlitro" in English number format. The remaining failures after the fixes are false flags (a title's "Phase 3" as "III fazy", "Two Year" as the compound "Dwuletnie", a group attribution across a "than" clause).
Planted number errors in translated sentences that both checks had passed (one digit changed): test 2,426 / 2,432 (99.8%; the preservation check alone 2,426, the cell trace alone 2,278), test2 917 / 920.
5. Speech
ASR, Qwen3-ASR-1.7B, FLEURS test, 100 sentences per language, language given; lower case, punctuation to spaces; digits left as written (a number spoken as a word counts as an error):
| de | fr | es | it | nl | pl |
|---|---|---|---|---|---|
| WER 15.2 / CER 4.97 | 4.75 / 1.90 | 2.36 / 0.91 | 8.60 / 2.93 | 6.78 / 2.21 | 14.01 / 4.82 |
German and Polish WER is inflated by compounds and numbers written differently from FLEURS's normalised text (e.g. "22 00 uhr" vs "22 uhr", "n chr" vs "nach christus"); the first 30 German items had only two over 15%. Voxtral was not running during this eval, so there is no head-to-head.
TTS intelligibility: 40 FLORES devtest sentences without digits per language, spoken by the house voice Ines (VoxCPM2, reference-clip mode) and transcribed by Qwen3-ASR; WER of transcript vs text (both models' errors count):
| WER | CER | WER with check-and-retry | re-rendered | |
|---|---|---|---|---|
| German | 23.35 | 14.41 | 14.19 | 14 / 40 |
| French | 15.85 | 9.39 | 10.13 | 10 / 40 |
| Spanish | 2.61 | 0.92 | 1.98 | 2 / 40 |
| Italian | 20.15 | 11.68 | 6.75 | 11 / 40 |
| Dutch | 6.73 | 1.87 | 5.38 | 7 / 40 |
| Polish | 5.14 | 1.36 | 4.77 | 4 / 40 |
The reference-clip mode keeps one voice across languages but sometimes garbles a German, French or Italian line (for example "Ring hat westabhaskuer..."); voice-design mode was cleaner on the few lines compared but changes the voice between calls. So the block transcribes every render and re-renders once when the transcript is more than 15% off; the second column is that loop. Continuation mode (reference + its transcript) was worse: it spoke "sześć" for a Polish "3".
The spoken check on sentences with numbers (20 test translations per language from the drift set: dosing, labels, specifications; one render each, no retry): flagged German 8/20, French 11/20, Spanish 11/20, Italian 8/20, Dutch 6/20, Polish 10/20. Reading the transcripts: almost all flags are real render errors, and they are serious. VoxCPM2 in this mode does not normalise symbols, so "5 V, 2 A", "78 dB(A)", "31/12/2026" and "EN 71-3:2019" come out garbled, and it sometimes invents content: a French "after 200 hours of use" was spoken as "every 15 to 20 years depending on the filter type", "plus de 100 kg" as "plus de 50 kilos", a Spanish "7 W" as "7000 lúmenes". Some flags come from the ASR side ("catorze" for 14). Verdict: do not use this TTS for safety text with numbers, units or dates without the spoken check and a person listening; spelling numbers out before synthesis (VoxCPM's normalize option, not tried) is the next step. That step is v2, section 7.
No native speaker has listened to any render, so no language is advertised for speech.
7. v2: numbers spelled out before speech (27 Sep 2026, branch the pre-release branch)
VoxCPM2's own normalize first. It runs wetext + inflect and handles English and Chinese only: on a German line it writes
"five V, two A", on a Polish one "one,five V". Measured anyway on the 20 v1 lines (the normaliser run on CPU with the voxcpm
package, which is exactly what normalize=True does before generation): 46 of 120 renders flagged, against 50 without it.
Not used.
The block's own normaliser (decosa_api/lang/spell.py, hand-written tables, no model, deterministic). Cardinals,
decimals with the language's separator word, signs, ranges ("220–240 V" -> "... bis ..."), scale words, units (SI and the
common ones: V, A, W, kWh, mAh, kg, mg, ml, °C, dB(A), h, min, mmol/l, mg/kg ...), numeric and month-name dates, times,
percentages, currency with cents and codes (digit runs of five or more read digit by digit), for en, de, fr, es, it, nl, pl.
Grammar where it changes the number word: gender for 1 (German ein/eine and dative/accusative after a preposition,
French un/une, Spanish un/una and feminine hundreds, Polish jeden/jedna and dwa/dwie), Polish 1 / 2-4 / 5+ noun forms,
Polish genitive and locative after prepositions, Polish ordinal dates and years ("trzydziestego pierwszego grudnia dwa
tysiące dwudziestego szóstego roku", "w dwa tysiące dwudziestym piątym roku"). Unit words the translator already wrote
("Stunden", "ans", "lat") are left alone. Each rewritten span keeps its source and spoken offsets.
The round trip reads the transcript back (digitise/readback): number words of any size to digits (including mixed Polish
cases such as "trzydzieści ośmiu"), ordinals in dates, decimals, "minus", unit words after a number to their symbols (only
for units the source has), clock times, amounts with cents, codes put back as written; then preserve.check compares it
with the source translation. num2words (LGPL-2.1) was used only as a cross-check while writing the tables and is not a
dependency. Self-consistency: the spelled-out text of 2,106 sentences (every drift source and translation, and every
sentence of the 5 lay-summary test2 trials in 7 languages) read back and checked against its source: 10 differ from the
source's own check (0.5%; compounds with a decimal inside such as "Ruxolitinib-1,5%-Creme", a 13-digit Dutch code, two
Spanish/French hour-only times).
Eval. The 20 v1 lines per language plus the next 20 test-split drift lines with a digit (held out: never looked at
while building), house voice Ines, VoxCPM2 reference-clip mode, seed 0, Qwen3-ASR-1.7B. Every transcript is scored by the
same final checker (read-back + preserve.check against the source). Real = the transcript lost or changed a number, date
or negation (judged by reading the transcript, no listening); artefact = the content is there and the read-back missed it.
Invented = a number, unit, date or time in the transcript that is not the source's. scripts/lang_eval_spoken_v2.py;
results language-pack/speech-spoken-v2.json, per render speech-spoken-v2-items.json.
| 240 renders (40 lines x 6 languages) | flagged | real errors | invented or changed figure |
|---|---|---|---|
| before: text as written | 94 (39%) | 88 (37%) | 52 (22%) |
| after: spelled out, one take | 46 (19%) | 44 (18%) | 23 (10%) |
| after: spelled out + the block's retry (one more take when flagged or > 15% WER) | 33 (14%) | 31 (13%) | 12 (5%) |
Held-out lines alone: before 44 / 120 flagged (43 real, 21 invented), after with retry 11 / 120 (11 real, 3 invented). The v1 check (no read-back) flagged 112 of the 240 "before" renders; the read-back removes its false flags on numbers the voice said as words. Mean WER against the spelled-out text: 0.23 before, 0.09 after with the retry.
Real errors per language, 40 lines each, before -> after with retry: German 12 -> 3, French 20 -> 11, Spanish 16 -> 4, Italian 15 -> 5, Dutch 12 -> 6, Polish 13 -> 2. What remains is the voice: in reference-clip mode VoxCPM2 sometimes speaks another sentence altogether ("Ah non, pas les talibans..." for a French weight limit), and standard codes ("EN 71-3:2019") are read wrong in most languages even spelled out. Verdict: no language is advertised for unattended spoken safety text; French fails outright (11 of 40 still wrong after the retry). The check flags these renders; how many bad renders it misses is not measured (that needs people listening).
8. v2: the meaning check (27 Sep 2026, branch the pre-release branch)
decosa_api/lang/meaning.py, POST /lang/meaning. Per segment: (1) Hy-MT2 translates the translation back into English
(the same :8491 service, a decosa.model-call.v1 receipt); (2) Qwen3.8-27B through the model gateway (receipted) gets the
source, the translation and the back-translation and answers in the typed-judgment style: ISSUE: type | source words | translation words lines from a closed set (negation, dropped, added, roles, entity, hedge, number), then VERDICT: OK|CHECK|ERROR and a reason. Code turns the quoted words into character spans (a quote that is not in the text gets no
span), drops issues that quote nothing from the source (grammar remarks) and keeps the model's verdict. Output per segment:
ok / check / error, issues with spans, reason, back-translation, receipt ids.
Planted set. 48 synthetic English sentences (product safety, medicine leaflets, lay summaries, notices; companies,
products and medicines invented; CC0), language-pack/meaning-sources.json. For each, hand-written minimal edits, one per
applicable type: negation flipped, clause or condition dropped, roles swapped, a named thing changed, a hedge changed. The
edited English is translated by Hy-MT2 and paired with the ORIGINAL source, so the planted translation is fluent. The
unedited translation is the clean pair. Pairs whose translation equals the clean one are dropped. Split by source before any
run: m00-m15 dev, m16-m47 test. Runner scripts/lang_eval_meaning.py; results meaning-summary.json, meaning-items.json.
Dev (tuning). Four prompt versions on dev (planted caught / clean pairs flagged): v1 271/279, 28/96; v2 274/279, 36/96; v3 271/279, 33/96; v4 (the model's OK verdict is kept even when it listed a difference) 267/279, 23/96 (1 as error). v4 is frozen.
Test (32 sources x 6 languages, run once, prompt frozen):
| planted | caught (check or error) | caught as error | |
|---|---|---|---|
| negation flipped | 106 | 102 | 100 |
| clause dropped | 144 | 144 | 136 |
| roles swapped | 90 | 89 | 85 |
| entity changed | 132 | 132 | 130 |
| hedge changed | 90 | 89 | 71 |
| all | 562 | 556 (98.9%) | 522 (92.9%) |
By language (caught / planted): de 93/94, fr 91/93, es 94/95, it 92/93, nl 94/94, pl 92/93. The 6 misses, read: all six are plants the translator had undone (Spanish and Polish "expose them to direct sunlight" translated as "avoid exposing them", German "stop taking suddenly" as "do not stop", French "should use ... unless" as "should use only if", a French role swap rewritten back, an Italian "could" rendered as "è possibile"); none is a planted error that reached the translation and was missed. Judged by the same agent that wrote the plants.
Clean pairs (test): 52 of 192 flagged (27%), 4 as error (2.1%). Read one by one: 7 are real problems in Hy-MT2's own translation (a Polish imperative turned into "they wear", "zlewek" = lab beakers for basins, an Italian sanding machine for a grinder, a Dutch "grondmachine", French "people outside" for bystanders, Dutch "leeg" = empty for a low battery, German "as soon as possible" for "as soon as you remember"); 5 debatable ("should" as "must" x3, a localised product name, "resident" for tenant); 40 false (mostly "can" for "may", near-synonyms, "in advance" made explicit). So a flag is a prompt to look, not a finding: about 1 clean line in 5 gets a CHECK, about 1 in 50 an ERROR.
Real lay-summary translations (lay summary test2: 5 real ClinicalTrials.gov trials, Hy-MT2 translations from the v1 eval):
- The known Polish "0 out of" drops: all 6 flagged ERROR ("W części 1 z 89 osób zmarło" -> "the translation omits the number 0, implying all 89 people died"). Of the 3 Polish "0 out of" sentences that kept the 0, 2 were flagged ERROR anyway (for the vaccine's name and "stan padaczkowy", both correct): false.
- A fixed random sample of 40 sentences per language (seed 66): flagged de 7, fr 11, es 8, it 8, nl 6, pl 14 (54 of 240; 14 as error). Read: 7 are real translation errors the number lock cannot see: "Vehicle Cream" (the inactive control cream) translated as a cream for cars in French, Spanish and Polish; "drop seizures" as "crisi ipertoniche" (Italian, a different seizure type); "Part 2 lasted up to 2; 72 mesi" (Italian); "the trial" as "Die Verhandlungen" (negotiations) and as "Proces" (a court case). 9 of the 14 ERRORs are false (e.g. "ensayo" read as "essay", "1 de cada 87" for "1 out of 87").
- "Decken Sie den Ladegerät nicht ab" (the GPSR charger run, wrong article): OK. Grammar, gender and articles are out of scope and the check says so.
Cost and latency. Per segment, test run: 627 prompt + 48 completion tokens on the gateway (Qwen3.8-27B), about $0.00026 at the typed-judgment list price ($0.30 / $1.50 per million), plus one Hy-MT2 back-translation on our GPU (about 100 tokens, 0.3 s). Wall time per segment p50 6.9 s, p95 10.1 s with 8 segments in flight on a shared, busy gateway; a 60-sentence lay summary in 6 languages is 360 segments, about 5 minutes at 8 in flight, which is why it is opt-in there.
Licence-clean QE models. Not used here; we then trained our own (MIT base, commercially usable data only): see
docs/evals/qe-model.md. As a first stage (/lang/meaning with qe: "cascade") it sends 14% of the planted test lines to this
LLM check, catches 530 of 562 (this check alone: 556) with 24 of 192 clean lines flagged (alone: 52). Earlier note: CometKiwi (Unbabel/wmt22-cometkiwi-da, wmt23-cometkiwi-da-xl) and XCOMET-XL are
CC BY-NC-SA 4.0 (non-commercial) and gated (Hugging Face model tags read 27 Sep 2026). The Apache-2.0 wmt22-comet-da (used in
section 1) needs a reference translation, which a live check does not have. A QE score would also not say which words
changed, which is the point here.
9. The term bank: per-customer, per-domain terminology memory (27 Sep 2026, branch the pre-release branch)
The meaning check's real catches (section 8) were domain terms: "Vehicle Cream" (the inactive control cream) as a cream for cars,
"the trial" as negotiations or a court case, "drop seizures" as a different seizure type. The number lock cannot see these, and
Hy-MT2 cannot know them without being told. decosa_api/lang/termbank.py, routes in termbank_routes.py.
What it is. An entry is a source term, per language the required renderings (any one must appear; a word ending in * takes
any ending, for inflecting languages), the forbidden renderings (the wrong sense), a domain tag, a note and provenance (a regulation,
a public glossary, a reviewer's fix, a meaning-check catch; who added it and when). Banks: one per API key (the customer), a
read-only demo bank, and public domain packs shipped with the code. Every change to a customer bank is a signed, append-only event
(decosa.termbank-event.v1, hash-chained per bank, Ed25519 by decosa-api's key, stored like the consent ledger in
<data>/lang/termbank.sqlite); a bank's state at version n is the fold of its first n events, so a customer can show which
version a translation used. The version reference (tenant bank id, seq and head hash, and each pack's sha256) goes into every
model call's params (term_bank_sha256) and into the signed translation receipt (term_bank).
Enforcement. Hy-MT2's model card documents a terminology prompt ("Reference the following translations: X translates to Y");
the existing glossary lock already used it, so the bank feeds its entries into the same lock (a customer's own request glossary
wins for the same term). After translation, code checks each source term: a required rendering must be present and no forbidden
one (new flag term_forbidden, high, with the target span). On a violation the segment is translated once more with the rule
spelled out ("Do not translate "Vehicle" as "pour véhicules""), in the model card's personalisation format; if it still fails it
stays flagged. A term inside a hyphenated compound ("anti-seizure") or a Title Case name ("the EU Clinical Trials website") does
not apply; English plurals do.
Learning loop. A meaning-check entity issue on a short phrase, or a reviewer's edit of a translation, becomes a proposal
(POST /lang/termbank/proposals): the wrong words as a forbidden rendering (only if they are really in the translation; the judge
sometimes quotes the back-translation), the required rendering from a public pack when one covers the term, else left for the
reviewer. Proposals never apply. Only an explicit accept by the key's owner turns one into an entry, and the accept event records
who and the final wording.
Public packs (decosa_api/lang/data/termpacks/, version 2026-09-27.1):
| Pack | Terms | Built from | Licence |
|---|---|---|---|
| clinical-trials | 12 (clinical trial, trial, placebo, vehicle, side effect, adverse event, serious adverse event, seizure, informed consent, ethics committee, investigational medicinal product, sponsor) | Regulation (EU) No 536/2014 Art. 2 in de, fr, es, it, nl, pl (EUR-Lex via the Publications Office Cellar); IATE entries cited per entry (1667431 placebo, 1804104 vehicle in pharmaceutical forms, 1108802 side effect, 1863780 adverse event, 1502886 seizure; wrong senses from 3561668 / 49476 / 3570563 court trial, 1240115 / 1368662 road and rail vehicle, 3584807 / 1722607 seizure of goods); three forbidden forms from the section 8 catches | EU legislation and IATE: reuse with acknowledgement under Commission Decision 2011/833/EU (IATE's licence as listed on data.europa.eu, publisher the Translation Centre; the EUR-Lex legal notice itself could not be fetched here) |
| product-safety | 16 (GPSR Art. 3 terms: manufacturer, authorised/authorized representative, importer, distributor, fulfilment service provider, economic operator, online marketplace, consumer, dangerous product, recall, withdrawal, market surveillance authority; IATE: warning, instructions for use, batch number) | Regulation (EU) 2023/988 Art. 3 in the six languages; IATE 1106944, 3535039, 841579, 152776, 950652, 3535591, 1679713, 1575564 | same |
Forms marked "editor" in an entry's sources are common usage added by the author (for example German "Importeur" beside the
GPSR's "Einführer", Spanish "evento adverso"). NCI Thesaurus is CC BY 4.0 (checked) but English-only, so it was not needed;
CDISC terminology was not used and its licence not checked; MedDRA is licensed and not used; EUPATI's glossary was not used. Not
covered: "treatment arm" (no source has the trial sense), "drop seizure" (no source found; the demo bank has it as a customer
entry), Portuguese and Czech.
The lay summary Member-State versions (/laysummary/translate) use the clinical-trials pack and the caller's bank by default;
the GPSR pack (/gpsr/pack) the product-safety pack; /lang/translate only when asked (term_bank).
9a. The seven real errors (lay summary test2, the section 8 sample)
The five test2 trials re-translated by TranslateRun exactly as /laysummary/translate does, in three configurations: no bank,
the clinical-trials pack, and the pack plus a customer bank holding one accepted proposal ("drop seizure" -> it "crisi con caduta",
the demo bank's entry). Two runs: the first as built, the second after two fixes the first run showed (below).
| Error (section 8) | No bank (run 1 / run 2) | Pack | Pack + accepted entry |
|---|---|---|---|
| fr "Vehicle Cream" -> "Crème pour véhicules" | reproduced / reproduced | "la crème placebo" | same |
| es "Vehicle Cream" -> "crema para vehículos" | reproduced / reproduced | "la crema con placebo" | same |
| pl "Vehicle Cream" -> "kremem do pojazdów" | reproduced / reproduced | "z kremem placebo" | same |
| de "The trial" -> "Die Verhandlungen" | reproduced / not reproduced ("Die Studie") | "Die Studie" | same |
| pl "The trial" -> "Proces" | reproduced / reproduced | "Badanie" | same |
| it "drop seizures" -> "crisi ipertoniche" | reproduced / reproduced | "crisi" (the wrong type gone, but "drop" lost) | "crisi con caduta" |
| it "72 months" -> "2; 72 mesi" (a garbled number, not a term) | not reproduced / not reproduced | - | - |
Prevented: 5 of the 6 term errors with the public pack alone, the sixth half-fixed; 6 of 6 with one accepted customer entry. The seventh error is not a terminology error and did not recur (vLLM batching makes greedy output vary slightly between runs; the German one also came out right in the second baseline run).
Side effects on the whole 5 trials x 6 languages (about 2,330 MT calls per configuration; the bank adds 8 retries):
- Run 1 (as built): the pack made one Dutch and one Polish sentence per trial fail: "EU Clinical Trials website" (a name) and "anti-seizure medicines" (a compound) were read as the terms. Fixed (names and compounds skipped, Dutch "klinische proeven" accepted), so these fixes are informed by this run.
- Run 2: failing sentences no bank -> bank: de 0 -> 0, fr 0 -> 0, es 0 -> 0, it 0 -> 0, nl 1 -> 0, pl 7 -> 6 (the six are the known Polish "0 out of" drops). "Check" sentences rose by 27, all one pattern: with the vehicle term translated, Hy-MT2 also wrote "BID" out as "twice a day", and the number lock flags the added number word. Those translations are right; the lock does not know BID.
- Meaning check on the 249-sentence sample (section 8's 240 plus the 9 Polish "0 out of" sentences): flagged no bank 58 (18 error), pack 64 (21), pack + entry 61 (18). The 11 sentences whose verdict got worse with the bank, read: 6 are the judge objecting to the pack's term itself ("vehicle cream" -> "Placebo-Creme" read as a different entity), 2 are real new errors (German "Vehicle Cream BID" -> "Placebo-Creme-Tagesdosis", which loses "twice"; Italian "150 people" -> "Circa 150"), 3 are judge noise.
9b. Planted domain terms (synthetic, 6 languages)
60 English sentences written for this eval (termbank-planted-sources.json, CC0): 30 clinical-trial lay-summary lines and 30
GPSR product-safety lines, each with at least one domain term whose everyday or legal sense differs (trial, vehicle, seizure,
recall, withdrawal, sponsor, authorised representative...). Split before any run: every third sentence dev (20), the rest test
(40). Each translated by Hy-MT2 with and without the domain's pack. The packs were frozen after the dev run (one change: Polish
"działań niepożądanych" accepted). Test, run once:
| Test: 240 segments, 435 term occurrences | No bank | With the pack |
|---|---|---|
| Terms rendered by the pack's rules | 315 (72.4%) | 434 (99.8%) |
| Forbidden (wrong-sense) renderings by rule | 9 | 0 |
| Model calls | 240 | 241 (one retry) |
By language (terms right by the pack's rules, no bank -> pack): de 52 -> 72 of 73, fr 61 -> 73, es 59 -> 73, it 47 -> 73, nl 48 -> 73, pl 48 -> 70 of 70. Dev: 128 -> 191 of 192.
The "with the pack" column is graded by the rules that steered it, so every term the bank changed was read (by the agent that
wrote the packs and the sentences; termbank-planted-test-read.json). Of the 120 term occurrences the pack's rules rejected
without the bank:
| Reading of the no-bank translation | Terms |
|---|---|
| Wrong sense (a court case, a car, proceedings; "withdrawal" and "recall" swapped) | 15 |
| Meaning shift across a defined distinction (adverse event rendered as side effect; recall rendered as the GPSR's withdrawal) | 14 |
| Non-standard or wrong term a reader can still decode (Spanish "servicios de cumplimiento" for fulfilment, "Wirtschaftssubjekt") | 17 |
| Qualifier dropped ("informed" consent) | 1 |
| An acceptable everyday synonym that is not the regulation's term ("patrocinador", "Ethikkomitee", "untersuchtes Arzneimittel") | 67 |
| The pack's forms too narrow; the translation was right (Polish "placebem", adjective-first word order, German verb forms) | 6 |
With the pack none of the 47 real problems (the first four rows) remains. The damage, read on all 240 translations with the pack: 8 have a grammar or spelling error the forced term caused ("Im Placebo-Gruppe", Italian "la medicinale sperimentale", Polish "Konsumenti", "produktzie") and 1 has a clause the translator added ("op verzoek van de veiligheidsinstanties"). The meaning check flagged 79 of 240 without the bank (21 as error) and 92 with it (29 as error); of the 48 whose verdict got worse, 30 are the judge objecting to the pack's term (it reads "placebo" for "vehicle", "mandataire", "Studie" for "trial" and "recuperación" as different things), 2 are real new errors (the added Dutch clause; French "chauffeurs" for heaters), 16 are judge noise unrelated to terms.
Verdict. The bank reliably removes wrong-sense domain terms: 15 wrong-sense and 14 meaning-shift renderings of 435 on the test set went to 0, and 6 of 6 real lay-summary term errors with one reviewed customer entry. It costs almost nothing (one retry in 240 segments here; the terminology prompt does the work on the first pass). The price is choice: the packs enforce the regulation's defined terms, so 67 acceptable lay synonyms were also replaced, and some forced terms break agreement (8 of 240). The meaning check does not yet know the bank and objects to its terms; feeding the applied terms into the judge's prompt is the next step. A native speaker has not reviewed any rendering; the packs cite their sources, not a linguist's sign-off.
Runner: scripts/lang_eval_termbank.py (planted, real, meaning, score). Results: termbank-planted-{dev,test}.json,
termbank-planted-test-read.json, termbank-real.json (run 2), termbank-real-asrun.json (run 1), termbank-meaning-*.jsonl,
termbank-summary.json. MT on the running decosa-lang-mt (no new GPU load); the meaning check through the gateway (about 1,500
judgments, about 0.9 M prompt and 70 k completion tokens, about $0.40 at list price).
10. All 24 EU official languages (27 Sep 2026, branch the pre-release branch)
Hy-MT2-7B's card lists 33 languages; nine of them are EU official languages (en, de, fr, es, it, nl, pl, pt, cs). The
other fifteen (sv, da, fi, el, ro, hu, bg, hr, sk, sl, lt, lv, et, ga, mt) are not on it, and Hy-MT2-30B-A3B has the
same list, so the larger model adds none. The block now routes each target language to a model
(decosa_api/lang/route.py, table decosa_api/lang/data/routing.json, written by scripts/lang_routes.py from
docs/evals/language-pack/eu-scores.json).
Candidates and licences (model cards read 27 Sep 2026; no non-commercial model considered):
| Model | Licence | Size | EU languages | Result |
|---|---|---|---|---|
Qwen3.8-27B (Qwen/Qwen3.8-27B @ 1d4bf0f2) |
Apache-2.0 | 27B, already served (the hosted model, qwen3.8-27b) |
all, not advertised per language | served for the fifteen |
| Hy-MT2-30B-A3B | Apache-2.0 | 30B MoE | the same nine as the 7B | no new language; not run |
EuroLLM-9B-Instruct-2512 (utter-project @ def82454) |
Apache-2.0 | 9.2B (18.3 GB bf16) | all 24 | measured; best on 14 of 15; proposed, not deployed (no GPU room) |
| EuroLLM-22B-Instruct-2512 | Apache-2.0 | 22.6B | all 24 | screened (6 languages x 200 sentences): level with the 9B (Maltese +1.5 COMET, others within 0.5) at 2.5x the memory; not worth it |
Salamandra-7B-instruct (BSC-LT @ a3ed5452) |
Apache-2.0 | 7.8B | all 24 | screened (15 x 200): under Qwen3.8-27B in 13 of 15 and under EuroLLM-9B in all 15 |
MADLAD-400-3B-MT (google @ fa184c67) |
Apache-2.0 | 2.9B | all 24 | screened on 100 sentences with greedy decoding: far behind (COMET 28-86, most under 55; English words left in); our decoding was not tuned (no beam search), so this is not a verdict on the 7B/10B models |
| TildeOpen-30B | CC-BY-4.0 | 30.7B, base model (no instruction tuning) | all 24 | not run: a base model at 30B does not fit the ceiling; Tilde's smaller models are WMT26 compression entries for cs-de only |
| NLLB-200, SalamandraTA-7B-instruct | CC-BY-NC-4.0, GPL-3.0 | - | - | excluded (licence) |
Method. FLORES-200 devtest (CC-BY-SA 4.0), all 1,012 sentences, English into each language; the block's own prompt
(Hy-MT2's) for Hy-MT2 and Qwen3.8-27B, greedy, Qwen with thinking off, through the same client the block uses (Qwen via
the model gateway: scripts/lang_eval_flores.py --backend gateway). Candidates in MLX on the Mac (bf16, greedy, their
model card's translation prompt: mlx_flores.py/madlad_flores.py, kept outside the repo with the per-sentence outputs).
chrF++ from sacrebleu 2.6.0 (nrefs:1|case:mixed|eff:yes|nc:6|nw:2); COMET-22 (Unbabel/wmt22-comet-da, Apache-2.0)
x100, scored on the Mac (MPS); it reproduces our server's CUDA score for Hy-MT2 German exactly (89.02). Qwen latency
at 8 requests in flight on the shared gateway (other evaluation jobs running): p50 1.6-3.7 s, p95 2.6-9.3 s per sentence.
The quality bar. A language is ready when its routed model scores COMET-22 >= 85.0 and chrF++ >= 45.0 and COMET can
judge the language; otherwise draft, and every translation into it comes back with tier: "draft" and
review: "draft, needs a reviewer". Why there: 85 sits 2.6 points under the weakest language the block already offered
(Spanish, 87.59), so "ready" means "in the band of what we already ship". The measured languages split cleanly: nothing
between 73.4 and 87.6, so where exactly the bar sits in that gap changes no route. COMET's encoder (XLM-R) was not
pretrained on Maltese, so a Maltese COMET score cannot vouch for it: Maltese stays draft whatever it scores. The chrF++
floor catches a model that COMET rates kindly but that loses surface form (Irish is under both). This is a routing rule,
not a quality guarantee: FLORES is Wikipedia-style prose, COMET is not calibrated across languages, and no native
speaker has read any language.
Routing rule. Per target language, the higher COMET-22 among the served models (Hy-MT2-7B only for its languages, with a 0.3-point preference for it: a dedicated 7B, no load on the shared LLM). Hy-MT2-7B is also only used when the source language is one of its own (Finnish into German goes to Qwen).
| Language | Route | chrF++ | COMET-22 | Qwen3.8-27B COMET / chrF++ | Hy-MT2-7B COMET / chrF++ | EuroLLM-9B COMET / chrF++ (not served) | Tier |
|---|---|---|---|---|---|---|---|
| Bulgarian (bg) | qwen3.8-27b | 61.70 | 91.25 | 91.25 / 61.70 | - | 91.70 / 64.42 | ready |
| Croatian (hr) | qwen3.8-27b | 56.07 | 91.05 | 91.05 / 56.07 | - | 90.53 / 54.56 | ready |
| Czech (cs) | hy-mt2-7b | 57.99 | 92.51 | 91.87 / 55.66 | 92.51 / 57.99 | 92.07 / 56.15 | ready |
| Danish (da) | qwen3.8-27b | 64.81 | 90.83 | 90.83 / 64.81 | - | 91.63 / 67.80 | ready |
| Dutch (nl) | hy-mt2-7b | 55.42 | 88.63 | - | 88.63 / 55.42 | 88.84 / 56.13 | ready |
| Estonian (et) | qwen3.8-27b | 52.73 | 90.29 | 90.29 / 52.73 | - | 92.07 / 57.10 | ready |
| Finnish (fi) | qwen3.8-27b | 52.98 | 91.67 | 91.67 / 52.98 | - | 92.74 / 55.00 | ready |
| French (fr) | hy-mt2-7b | 67.47 | 89.04 | - | 89.04 / 67.47 | 89.00 / 69.55 | ready |
| German (de) | hy-mt2-7b | 62.98 | 89.02 | 88.69 / 62.99 | 89.02 / 62.98 | 88.82 / 64.37 | ready |
| Greek (el) | qwen3.8-27b | 50.02 | 89.25 | 89.25 / 50.02 | - | 90.03 / 51.63 | ready |
| Hungarian (hu) | qwen3.8-27b | 52.39 | 89.37 | 89.37 / 52.39 | - | 90.21 / 54.02 | ready |
| Irish (ga) | qwen3.8-27b | 44.60 | 73.41 | 73.41 / 44.60 | - | 81.26 / 54.06 | draft (COMET-22 73.41 < 85.0; chrF++ 44.6 < 45.0) |
| Italian (it) | hy-mt2-7b | 56.53 | 89.64 | - | 89.64 / 56.53 | 89.41 / 58.54 | ready |
| Latvian (lv) | qwen3.8-27b | 53.72 | 89.13 | 89.13 / 53.72 | - | 90.96 / 57.65 | ready |
| Lithuanian (lt) | qwen3.8-27b | 53.23 | 90.12 | 90.12 / 53.23 | - | 91.13 / 55.24 | ready |
| Maltese (mt) | qwen3.8-27b | 55.07 | 68.95 | 68.95 / 55.07 | - | 71.66 / 64.45 | draft (COMET-22 68.95 < 85.0; COMET-22 does not cover Maltese (not in XLM-R's pretraining)) |
| Polish (pl) | hy-mt2-7b | 50.10 | 90.54 | - | 90.54 / 50.10 | 90.49 / 50.56 | ready |
| Portuguese (pt) | hy-mt2-7b | 69.38 | 90.60 | 90.30 / 68.30 | 90.60 / 69.38 | 90.09 / 68.86 | ready |
| Romanian (ro) | qwen3.8-27b | 61.81 | 91.05 | 91.05 / 61.81 | - | 91.56 / 64.26 | ready |
| Slovak (sk) | qwen3.8-27b | 55.25 | 90.54 | 90.54 / 55.25 | - | 91.30 / 58.12 | ready |
| Slovenian (sl) | qwen3.8-27b | 53.35 | 89.58 | 89.58 / 53.35 | - | 90.45 / 55.70 | ready |
| Spanish (es) | hy-mt2-7b | 55.19 | 87.59 | - | 87.59 / 55.19 | 87.34 / 55.13 | ready |
| Swedish (sv) | qwen3.8-27b | 64.65 | 91.07 | 91.07 / 64.65 | - | 91.66 / 67.30 | ready |
Hy-MT2-7B beats Qwen3.8-27B wherever both were run (Czech +0.64, German +0.33, Portuguese +0.30; French and Spanish on 23 Sep: +0.15, +0.35), so the eight stay on Hy-MT2. Italian, Dutch and Polish were not run on Qwen (the run was stopped to free the shared gateway; Hy-MT2's own scores clear the bar there). 21 of 23 target languages are ready; Irish and Maltese are drafts.
Receipts. Every Hy-MT2 call keeps its decosa.model-call.v1 receipt; every Qwen call goes through the app's LLM
client and carries the gateway-signed receipt (checked end to end on the pre-release server: GET /receipts/chatcmpl-...
returns status: signed, the hosted model as upstream_model). The signed decosa.lang-translation.v1 receipt now
names each target's model and tier and lists the models used (repo, revision, licence, route); /lang/verify checks
it as before.
10.1 The number lock in the added languages
Word tables for all fifteen (units in each script, month names, number words, scale words, negations, "both"/"none"/
"without" words, ordinals). Code changes in preserve.py, all additive: Greek, Cyrillic and Romanian ș/ț count as
letters; year-first dates (Hungarian "2026. március 12.", Lithuanian "2026 m. kovo 12 d.", Latvian "2026. gada 12.
martā") and spaced numeric dates ("12. 3. 2026"); "klo 14.30" times; Romanian "20 de ore"; Maltese article hyphens
("sat-12", "fil-14:30"); ordinal units ("var 8:e timme", "hver 8. time"); Finnish and Hungarian case endings on figures
and codes ("500 mg-ot", "25 °C:ssa", "23-4471-es"); per-hour idioms ("litraa tunnissa", "λίτρα την ώρα", "литра в час");
Bulgarian "г." as a year after a calendar year and a gram otherwise. Irish and Maltese read 1,500.5 the English way.
Negations are not read for Slovak, Lithuanian and Latvian (a ne- prefix on the verb, as Czech) or Maltese (a -x suffix).
The drift eval (section 2's 90 synthetic sentences and planting code, with planting rules added for the five languages)
through the routed path (Qwen3.8-27B) into Finnish, Hungarian, Greek, Bulgarian and Romanian. The tables were written
against my own test sentences first, then fixed on the 30-sentence dev split (487/488 caught, 3 of 150 clean translations flagged);
the 60-sentence test split was run once with the checker frozen (drift-eu-test-frozen.json):
| Error planted | fi | hu | el | bg | ro | all |
|---|---|---|---|---|---|---|
| a digit changed | 51/52 | 52/53 | 53/53 | 53/53 | 53/53 | 262/264 |
| a number x10 | 46/46 | 47/48 | 49/49 | 49/49 | 49/49 | 240/241 |
| a thousands group written the English way | 2/2 | 0/0 | 2/2 | 0/0 | 2/2 | 6/6 |
| a unit swapped | 33/34 | 32/33 | 29/30 | 28/29 | 37/38 | 159/164 |
| a number and its unit dropped | 50/52 | 53/53 | 53/53 | 53/53 | 53/53 | 262/264 |
| the negation removed | 20/20 | 18/18 | 19/20 | 18/19 | 16/18 | 91/95 |
| a date changed | 6/6 | 7/7 | 6/6 | 7/7 | 6/6 | 32/32 |
| a code changed | 2/2 | 3/3 | 3/3 | 2/2 | 2/2 | 12/12 |
| all | 210/214 | 212/215 | 214/216 | 210/212 | 218/221 | 1,064 / 1,078 (98.7%) |
False flags on the clean test translations: 20 of 300 when frozen. One is real: Hungarian "under 36 months" came out "3 hónaposnál kisebb" (under 3 months). 19 are false: unit words the tables lacked (inches in five languages, the Greek "γρ.", "hétig"), the Bulgarian "г." read as a year, "Tiltott" (prohibited) missing from the Hungarian negations, the Finnish "1:20" for "1 in 20" read as a time, a word between number and unit ("200 käyttö tunnin"), and negation carried by a ne- prefix in Bulgarian and Romanian ("Неподходящ", "Nepotrivit"). After fixes made from those (so no longer held out): 1,068 / 1,078 caught, 9 of 300 clean flagged (the one real, eight false of the last three kinds).
Misses (frozen, 14): four were negation plants that left another negative word in place (Greek, Bulgarian and Romanian "no participant died" keep "Κανένας", "Нито един", "Niciun" once "δεν"/"не"/"nu" is gone; a Romanian "nicio" left in a garbled sentence), three are section 2's known case ("children aged 8 to 14" has no unit in English, the plant changed the unit the translator added), two a Hungarian code with a case ending and two a Finnish °C range read as a compound (both fixed after), two a day count where the ordinal "day 14" matched the changed "14 weeks", and one a plant in a sentence the checker had already flagged (Finnish "1:20").
The other ten added languages have tables, unit tests (tests/test_lang_eu.py: a clean dosing text passes and a planted
change is caught in all fifteen) and no planted-error run yet.
10.2 Lay summary Member-State versions in added languages
The five fresh trials of the lay summary eval (test2: real ClinicalTrials.gov results, 330 sentences) translated by
TranslateRun through the router into Finnish, Greek, Hungarian, Romanian and Irish (all on Qwen3.8-27B;
scripts/lang_eval_laysummary.py run --split test2 --langs fi,el,hu,ro,ga --tag=-eu). The tracer reads Irish with a
decimal point, the others with a decimal comma. Frozen code, then the stored translations re-checked after the fixes
below (recheck, not held out):
| numbers traced | wrong | sentences fail (frozen -> after fixes) | |
|---|---|---|---|
| Finnish | 322 / 326 | 0 | 30 -> 7 |
| Greek | 337 / 341 | 0 | 1 -> 0 |
| Hungarian | 342 / 345 | 0 | 2 -> 1 |
| Romanian | 304 / 308 | 0 | 0 -> 0 |
| Irish (draft) | 299 / 304 | 1 | 11 -> 6 |
Most Finnish failures were the checker: Qwen writes "86 out of 87" as "86/87", which the frozen code read as one number and a stray "/87". Fixes: a count/total fraction is read as two numbers; Hungarian month endings ("2017 novemberében"); Finnish case forms of 2 to 10 ("kolmeen ryhmään"); the Greek "του" between month and year; Irish copula negations ("Níorbh"); the Irish "sé" (six, and "he") no longer counts. The six-language results in section 2 and section 4 (test split) are unchanged by these fixes (re-run: 1,320 / 1,328 and the same recheck totals).
Real translation errors the checks found: Finnish "391 of the 631 people (62%) were women" lost the 391; Irish turned "November 2017" and "May 2024" into Nollaig (December) and Meitheamh (June), dropped "HbA1c" from a sentence it garbled, and mistranslated the trial title. Remaining failures after the fixes are the English abbreviation "BID" translated as "twice daily" (a number word the English side does not have: false), "12 to under 18" written "12-17" (a liberty), a Hungarian number said twice, and "COVID-19" rendered as "coronavirus pandemic" (a locked term). No native speaker read these versions.
10.3 Candidates on the added languages
Full devtest, EuroLLM-9B-Instruct-2512 against the route (COMET-22 / chrF++), from the table above: it is higher in 14 of the 15 languages Qwen serves (Croatian 0.52 lower), by +0.45 (Bulgarian) to +1.83 (Latvian) COMET and +1.6 to +4.4 chrF++ in the other twelve ready ones (Croatian is lower on both), and by far where Qwen is weak: Irish 81.26 vs 73.41 COMET (54.06 vs 44.60 chrF++) and Maltese 64.45 vs 55.07 chrF++. Irish would still be under the bar with it.
Same first 200 sentences, three candidates (COMET-22 / chrF++):
| Qwen3.8-27B (route) | EuroLLM-9B | EuroLLM-22B | Salamandra-7B | |
|---|---|---|---|---|
| Bulgarian | 91.22 / 61.3 | 91.55 / 64.2 | - | 89.91 / 59.7 |
| Croatian | 90.72 / 56.8 | 90.15 / 55.5 | - | 87.79 / 51.2 |
| Danish | 90.46 / 63.5 | 91.13 / 66.2 | - | 90.40 / 62.5 |
| Estonian | 90.18 / 53.3 | 91.81 / 57.2 | - | 89.45 / 51.4 |
| Finnish | 90.94 / 52.4 | 92.32 / 53.4 | 92.27 / 54.7 | 90.73 / 50.7 |
| Greek | 89.27 / 51.1 | 90.41 / 53.2 | 90.36 / 53.5 | 88.35 / 49.6 |
| Hungarian | 88.71 / 52.7 | 90.12 / 53.3 | 90.57 / 53.5 | 88.33 / 49.6 |
| Irish | 73.56 / 44.6 | 81.84 / 55.9 | 81.87 / 56.8 | 76.45 / 48.2 |
| Latvian | 88.56 / 52.8 | 90.55 / 57.0 | - | 86.73 / 48.7 |
| Lithuanian | 90.02 / 53.1 | 91.09 / 54.4 | 91.14 / 55.8 | 88.30 / 51.4 |
| Maltese | 68.33 / 55.5 | 70.98 / 65.5 | 72.44 / 68.5 | 68.54 / 59.9 |
| Romanian | 90.70 / 63.3 | 91.93 / 65.6 | - | 89.93 / 60.6 |
| Slovak | 90.64 / 54.1 | 91.17 / 57.8 | - | 89.47 / 52.9 |
| Slovenian | 89.27 / 53.9 | 90.40 / 56.5 | - | 88.53 / 53.2 |
| Swedish | 90.37 / 63.3 | 91.24 / 65.6 | - | 90.11 / 61.5 |
EuroLLM-9B on the languages Hy-MT2-7B serves (full devtest, COMET-22 / chrF++ against Hy-MT2-7B): German 88.82 / 64.37 (89.02 / 62.98), French 89.00 / 69.55 (89.04 / 67.47), Spanish 87.34 / 55.13 (87.59 / 55.19), Italian 89.41 / 58.54 (89.64 / 56.53), Dutch 88.84 / 56.13 (88.63 / 55.42), Polish 90.49 / 50.56 (90.54 / 50.10), Portuguese 90.09 / 68.86 (90.60 / 69.38), Czech 92.07 / 56.15 (92.51 / 57.99). It is about level with Hy-MT2-7B there (-0.51 to +0.21 COMET, mean -0.19; chrF++ higher in five of eight), so one EuroLLM-9B could serve all 24 languages in the slot Hy-MT2-7B holds today (in FP8; bf16 weights alone are 18.3 GB, which does not leave KV room at the current 0.18 setting).
Proposal (not deployed): replace Hy-MT2-7B on decosa-lang-mt with EuroLLM-9B-Instruct-2512 in FP8 at the same
memory, route every language to it (Qwen3.8-27B stays the fallback for Croatian if the numbers hold), after two checks this
branch did not do: FP8 quality on FLORES against the bf16 numbers here, and the drift set through it. Or, with about
12 GB more room on GPU0, run it beside Hy-MT2-7B for the fifteen. Either change is The owner's call: it swaps or adds a
resident model on the shared GPU.
VRAM. Nothing new is loaded: Hy-MT2-7B stays at 18.3 GB and the language pack at 29.7 GB of its 30 GB ceiling; the fifteen new languages use the Qwen3.8-27B already served on GPU1 through the gateway. GPU0 had 8.6-19 GB free during the day (measured with nvidia-smi, 78.6-89.3 GB of 97.9 GB in use). EuroLLM-9B-Instruct-2512 would need about 12 GB in FP8 (9.2 GB of weights plus KV at a 4k context; an estimate, not measured on the card), or about 20 GB in bf16. There is no room under the language pack's ceiling, so it is proposed, not deployed.
11. EU term packs, offline IATE and the term-aware meaning check (27 Sep 2026, branch the pre-release branch)
11a. Packs from EU legislation
Every EU regulation defines its terms in one article, numbered identically in all 24 language versions. scripts/eu_termpacks.py
fetches each act's Official Journal text as Formex 4 XML from the Publications Office Cellar by CELEX number (content negotiation:
Accept: application/zip;mtype=fmx4, Accept-Language: deu ...), caches the ZIPs with a sha256 manifest under
<internal path>, and decosa_api/lang/eulex.py reads the definitions article:
- The defined term is the quoted span (QUOT.START/QUOT.END) that opens each top-level point; a Spanish-style definition list (DLIST.ITEM with TERM) is read the same way. Versions without quote markup are read from italics (Swedish), typed quotes (some Finnish), or "term – definition" / "term: definition" (Lithuanian). Irish puts the verb first ("ciallaíonn ‘X’").
- Alignment is by the position of the point, checked three ways: the item count must match English, numeric labels must match,
and the numbers each definition cites (articles, acts) must agree on at least 60% of the points that cite any. A language that
fails is left out of the pack and named in
not_covered, never shifted. - "‘X’ or ‘Y’" and "‘X’ (‘Y’)" give alternative entries; a term defined in passing in parentheses ("(‘data subject’)") becomes its own entry when every language has the same count.
- Forms: the exact term plus a stem pattern from fixed suffix rules ("klinische Prüfung", "klinisch* Prüfung*"). Finnish and Czech versions name the term in an oblique case ("kliinisellä tutkimuksella", "klinickou studií"); their prompt form is a reliable IATE term that fits the stem pattern when there is one ("kliininen tutkimus", "klinická studie"). Hungarian leading articles are dropped and the Bulgarian disambiguating accent (ѝ) folded. A language term with a lead-in clause (much longer than the English, or with a stray number) is left out: 6 of 12,503.
- Deterministic: the same ZIPs give byte-identical packs (checked by building twice). Corrigenda and consolidated versions are not applied.
| Pack | Act | Where | Entries | Languages |
|---|---|---|---|---|
| eu-ctr | Regulation (EU) No 536/2014, Clinical Trials | Art. 2 | 35 | 24 |
| eu-gpsr | Regulation (EU) 2023/988, GPSR | Art. 3 | 28 | 24 |
| eu-mdr | Regulation (EU) 2017/745, MDR | Art. 2 | 73 | 24 |
| eu-ivdr | Regulation (EU) 2017/746, IVDR | Art. 2 | 76 | 24 |
| eu-ai-act | Regulation (EU) 2024/1689, AI Act | Art. 3 | 68 | 24 |
| eu-dsa | Regulation (EU) 2022/2065, DSA | Art. 3 | 24 | 24 |
| eu-dma | Regulation (EU) 2022/1925, DMA | Art. 2 | 33 | 24 |
| eu-gdpr | Regulation (EU) 2016/679, GDPR | Art. 4 | 27 | 24 |
| eu-esrs | Delegated Regulation (EU) 2023/2772, ESRS | Annex II, Table 2 | 199 | 22 (Italian and Lithuanian sort the glossary in their own alphabetical order, so row alignment fails and they are left out) |
563 entries, 12,497 language renderings. CSRD (Directive (EU) 2022/2464) was not built: its definitions are amendments inserted into Article 2 of Directive 2013/34/EU, not a definitions article of its own; the ESRS glossary carries the reporting terms.
Reuse. EUR-Lex legal notice, read from the page on 27 Sep 2026: "The Commission's document reuse policy is based on Decision 2011/833/EU. Unless otherwise specified, you can re-use the legal documents published in EUR-Lex for commercial or non-commercial purposes." Decision 2011/833/EU (read from the Cellar) Art. 4 makes documents reusable for commercial or non-commercial purposes and Art. 6(2)(a) allows the condition that the reuser acknowledges the source. Every pack carries the acknowledgement and each entry its act, point, ELI and CELEX. (Section 9's two hand-made packs said the legal notice could not be fetched; it has now been read, and their licence line is updated, version 2026-09-27.2.)
Alignment spot-check (eu-termpacks-spotcheck.json): seed 2772, 8 entries per pack (main terms, not alternatives), 6
languages each rotating over the 23, so every language is read; 428 readings of the language term against the English term,
with the source text read where unsure. First build: 0 wrong terms, 3 partial (Hungarian terms that kept the article "a/az");
after the article rule, 0 of 428 wrong or partial (a 95% upper bound of about 0.7% by the rule of three). One form issue (the
Bulgarian "ѝ") was folded. Read by the agent that wrote the extractor, not by native speakers.
Independent check against IATE (eu-termpacks-summary.json): of the 12,497 renderings, 5,472 have the English term in the
offline IATE index; for 4,302 of those (78.6%) a reliable IATE term matches the pack's forms. The rest are mostly IATE entries for
another sense of the same English word or older wording (the export is from 2019, before the AI Act, DSA, DMA, GPSR and ESRS); they
were not read one by one.
11b. Offline IATE
The official full download (linked from iate.europa.eu/download-iate; https://iate.europa.eu/em-api/artifacts/full-tbx) is
IATE_export_26022019.tbx: the export has not been refreshed since 26 Feb 2019. ZIP 112.8 MB (sha256 ac50c2a7...), TBX 1.86 GB,
934,921 entries. Licence: the European Commission reuse notice (Decision 2011/833/EU), as the IATE dataset's distributions on
data.europa.eu state (read through the data.europa.eu API on 27 Sep 2026; publisher: the Translation Centre). IATE's own
legal-notice page is rendered by JavaScript and was not read.
scripts/iate_build.py keeps entries whose subject field (EuroVoc codes) is in our domains: health (2841), law (12xx, 10xx EU law),
trade (20xx), product safety (6411 technical regulations, 2026 consumption), finance (24xx, 4026 accounting), information
(3231, 3236; added for the AI Act, DSA and GDPR packs). Result: 351,505 entries, 2,701,051 terms in 24 languages, SQLite 401 MB at
<internal path> (outside git). Reliable terms (codes 3-4) by language range from 12,559
(Croatian) to 241,145 (French); the newer languages have far fewer (Czech 20,827, Bulgarian 22,187).
Use: suggestions only. Proposals whose required rendering is unknown carry suggestions (reliable, not deprecated IATE terms,
domain-filtered), and GET /lang/termbank/iate answers a reviewer. Nothing from IATE is enforced unless a person accepts it.
Without the index, suggestions are simply absent.
11c. The term-aware meaning check
Section 9 found the judge objecting to the pack's own terms ("placebo" for "vehicle"). Now each segment's terms whose required
rendering is in the translation, and whose forbidden rendering is not, are listed to the judge as REQUIRED TERMS with a rule that a
listed rendering is correct by definition and that every other word is judged as strictly as before; code then drops an entity issue
that is only about such a term (the rest of its quote is an article or is still in the back-translation), and records it in
term_suppressed. Segments without an applicable term get the unchanged v4 prompt.
Method: only segments with at least one applicable term are re-judged (1,006 across the sets), each twice on the same day and gateway: v4 (before) and v4 + terms (after). Translations reused from sections 8 and 9.
Tuning, disclosed: two changes on planted-dev (a term whose forbidden rendering is present is not listed, after run 1 cleared two
real car-gel errors; only entity issues are dropped and the exemption is limited to the listed words, after run 2 cleared
"épilepsies" for seizures). After the test sets ran, one more change: the as-run rule dropped entity issues whose quote was the
term plus up to two words, which cleared three planted errors ("Serious side effects" as "mild side effects", fr/nl/pl) and masked
the real Italian "drop seizures" error. The final rule (article, or a word the back-translation kept) was re-applied to the same
runs; the prompt did not change, so the flagged/not-flagged numbers below are exact for it. As-run numbers are in
meaning-terms-summary.json.
| Segments with applicable terms | n | Flagged before | Flagged after | Flagged only for a term objection, before -> after |
|---|---|---|---|---|
| 9b planted test, with the pack | 238 | 90 | 57 | 51 -> 0 |
| 9b planted test, no bank | 199 | 59 | 66 | 16 -> 0 |
| 9a real lay summaries, pack | 96 | 30 | 19 | 13 -> 0 |
| 9a real, pack + accepted entry | 96 | 28 | 28 | 11 -> 0 |
| 9a real, no bank | 81 | 22 | 23 | 4 -> 0 |
| Section 8 planted errors (test) | 54 | 54 caught | 54 caught | - |
| Section 8 clean pairs (test) | 24 | 5 | 5 | 1 -> 0 |
Real catches kept (meaning-terms-catches.json): Spanish "crema para vehículos" (no bank) error -> error; Italian "crisi
ipertoniche" for drop seizures error -> check; Italian "Circa 150" error -> error; German "Placebo-Creme-Tagesdosis" (loses
"twice") error -> check; French "chauffeurs" for heaters error -> error; the Dutch added clause error -> error; French
"épilepsies" (dev) error -> error; all 54 planted meaning errors that had a pack term. The other section 8 catches (French and
Polish car creams, Polish "Proces", German "Verhandlungen", the Polish "0 out of" drops) have no applicable term in those
translations, so they get the unchanged prompt.
Reading the changes: objections to pack terms are gone. The total falls with the pack (90 -> 57, 30 -> 19) but not everywhere, because the "as strictly as before" wording makes the judge flag more of the other words: on the planted no-bank translations 44 segments were newly flagged across the test set; of the first 30 read, 4 are real (French "acceptation éclairée" for informed consent, German "Versuch" and "Nebenereignis", Spanish "retirada" for recall) and the rest are noise of the section 8 kind (tense, "16 Jahren", singular and plural). A flag is still a prompt to look.
Cost: listing terms adds about 200 prompt tokens per judged segment that has terms (planted test: 624 -> 819 per segment); the eval used about 2.1 M prompt and 0.15 M completion tokens on the gateway (about $0.85 at list price) and back-translations on the running decosa-lang-mt.
11d. Wiring
GET /lang/termbank/packslists every pack with its languages, per-language counts, source document (CELEX, ELI, EUR-Lex link, ZIP hashes) and use cases, plus the menu per use case: lay summary defaultclinical-trials, optionaleu-ctr; GPSR defaultproduct-safety, optionaleu-gpsr,eu-mdr,eu-ivdr; the language pack any pack, none by default.term_bank: {"add": ["eu-mdr", "eu-ivdr"]}adds packs to a route's defaults (a device sold to consumers: GPSR plus MDR).- Languages without an entry degrade quietly: no prompt line and no flag for that term there, and the run's
term_bank.coveragenames the terms left unenforced per language. The hand-made packs cover 6 languages; the EU packs 22-24.
Runners: scripts/eu_termpacks.py fetch|build|sample|score, scripts/iate_build.py fetch|build,
scripts/lang_eval_meaning_terms.py run|score|catches. Results: eu-termpacks-build.json, eu-termpacks-spotcheck.json,
eu-termpacks-summary.json, meaning-terms-*.jsonl, meaning-terms-summary.json, meaning-terms-catches.json.
6. Receipts
Every MT, TTS and ASR call yields the generic decosa.model-call.v1 receipt (decosa_api/model_call.py, merged from
the pre-release branch): model id, weights (repo, revision, hf-files-v1 root over every file at the revision, computed
with scripts/model_root.py from the Hub), runtime and device, input and output part hashes, parameters (prompt id and
hash, decoding), units (characters and tokens for MT, audio_ms for speech), timings, signed by decosa-api's Ed25519 key,
stored and served at GET /receipts/{id}; with the gateway witness on it is also reported to the model gateway. These are
operator attestations, not proofs. Each translation run also gets a signed decosa.lang-translation.v1 over source,
targets, glossary and the model-call receipt ids. The eval runners ran before the merge and used the interim statement
(the first version of the block); the receipts do not change any number above.
Checkable properties of the sample runs (rehearsal/language-pack)
- The two-line dosing text translates into German and French with no high flag.
- A German translation with 500 mg written as 50 mg fails with
number_changedon "500 mg". - The same translation with "do not" dropped gets
negation_missing. - The translation receipt verifies against the source text.
/lang/checkneeds no token and makes no model call.- Finnish goes to
qwen3.8-27b, German stays onhy-mt2-7b, and Maltese comes back withreview: "draft, needs a reviewer".
Limits
The QE model (
/lang/qe,docs/evals/qe-model.md) is a fast screen, not a substitute for this check: it misses wrong-sense words (0 of 9 real ones) and more role and entity swaps than the LLM.The term bank (section 9) enforces the packs' defined terms, including where an everyday synonym would do; forced terms sometimes break agreement; no linguist has reviewed the packs. The EU packs (section 11) are generated: their forms are the regulation's exact term plus a rule-made stem pattern, their alignment was spot-checked by the builder (0 of 428 wrong), and generic defined words ("product", "risk", "consumer") are enforced wherever they occur once a pack is chosen.
IATE suggestions come from the official 2019 export; terms coined since (AI Act, DSA, GPSR) are not in it.
Meaning is checked by a model (section 8): it catches planted negation, clause, role, entity and hedge changes well but also flags about a fifth of correct lines for a look; grammar, gender and articles are not checked; no native-speaker review yet.
Spoken safety text: not advertised in any language (section 7); French fails outright.
All 24 EU official languages are served (section 10). Irish and Maltese are drafts; FLORES is Wikipedia-style prose, so the bar says nothing about legal or medical register, and COMET-22 cannot judge Maltese. The meaning check, spoken check and term packs were measured on the first six languages only; the term packs cover six, so the other languages get no term enforcement (and no error).
Negations are not read for Czech, Slovak, Lithuanian, Latvian (verb prefixes) or Maltese (a -x suffix); a negation carried by a prefix on an adjective ("Nepotrivit", "Неподходящ") is a false flag in the other languages.
The number lock in the added languages was tested with planted errors in five of fifteen (fi, hu, el, bg, ro); the other ten have word tables and unit tests only. The candidate models ran on a Mac in MLX, not in the served vLLM build.
In the text checker, number words above 20 (except tens and 100), ordinals beyond third, Roman numerals and number compounds ("Dwuletnie") are not read and produce false flags (the spoken check's read-back does read them).
The same author wrote the synthetic sentences, the planting code, the word tables and the checker, and judged the false flags.