Language pack
Translation into all 24 EU official languages with a number lock, each language served by the model that measured best for it. Every number, unit, date, time, code, negation and required term in a source line must come through its translation, read in that language's own number format; anything that drifts is flagged with its source span, after one constrained retry. Plus speech in consented house voices, speech recognition in the same languages, and a signed receipt for every model call. Open weights, Apache-2.0, 30 GB of GPU in total; the fifteen languages added on 27 Sep run on the Qwen3.8-27B we already serve, with no new GPU.
Measured 2026-09-27. Full eval. Used by GPSR listing pack and EU trial lay summary (Member-State versions).
What the lock catches
A German translation of a dosing line, checked in code with no model call (POST /lang/check):
English
Do not take more than 8 tablets (500 mg each) in 24 hours.
German (bad translation)
Nehmen Sie mehr als 8 Tabletten (je 50 mg) in 24 Stunden ein.
- Number changed: "500 mg" became "50 mg"
- Negation lost: the source says "Do not" and the translation has no negation: check the meaning did not flip
How it works
- Hy-MT2-7B (Apache-2.0, Tencent) translates into the eight EU languages it was trained on; Qwen3.8-27B (Apache-2.0), already served through our gateway, translates into the other fifteen. The route per language comes from measured FLORES-200 scores; languages under the quality bar come back marked "draft, needs a reviewer". Same prompt, greedy decoding, a terminology prompt when a customer glossary applies.
- Code reads numbers in each language's format (3,000.5 in English, 3.000,5 or 3 000,5 in German, French and Polish), units in symbols and words (8 hours = 8 Stunden = 8 godzin), scale words, currencies, ranges, dates with month names in each language, times, codes, e-mail addresses, URLs, phone numbers and negations.
- It matches the source's mentions to the translation's and flags what is left: a changed or missing number, a changed or lost unit, a changed date, a changed code, a lost 'not', a missing locked term. A German translation that keeps the English '3,000 mg' reads as 3 mg and is flagged.
- A line with a high-severity flag is translated once more with the numbers and terms spelled out, and the better attempt is kept; what still fails stays in, flagged. Nothing is rewritten silently by a second model.
- Meaning (POST /lang/meaning): Hy-MT2 translates each line back into English and Qwen3.8-27B compares it with the source for flipped negations, dropped clauses, swapped roles, wrong names and changed hedges, with the words it means marked in both texts.
- Term bank: per customer and per domain, required and forbidden renderings of domain terms (a 'vehicle cream' is a placebo, not a car cream), fed into the terminology prompt and checked after translation, versioned in signed receipts. Public packs for clinical-trial lay summaries and GPSR product safety, and nine generated from the definitions articles of EU legislation (CTR, GPSR, MDR, IVDR, AI Act, DSA, DMA, GDPR, ESRS) in up to 24 languages. The meaning check is told the bank's required terms.
- Speech: VoxCPM2 speaks in two house voices designed from a text description, each with a revocable consent-ledger entry; the render is transcribed by Qwen3-ASR-1.7B and its numbers are checked (a take that drifts is rendered once more).
- Every Hy-MT2, TTS and ASR call carries a model-call attestation (model revision, a hash over the weight files, input and output hashes, parameters, signed by decosa-api); every Qwen call carries the gateway's signed receipt. Each translation gets a signed receipt over source, targets and glossary that names the model and tier of each language.
All 24 EU languages: model and score per language
FLORES-200 devtest (CC-BY-SA 4.0), all 1,012 sentences, English into each of the 23 other EU official languages, the block's own prompt, greedy; chrF++ from sacrebleu 2.6.0, COMET-22 (Unbabel/wmt22-comet-da, Apache-2.0) x100. Each language goes to the model that scored higher: Hy-MT2-7B for the eight it was trained on, Qwen3.8-27B (already served, no new GPU) for the other fifteen. 21 of 23 meet the bar; Irish and Maltese are returned marked "draft, needs a reviewer".
| Language | Model | chrF++ | COMET-22 | Other model's COMET | Status |
|---|---|---|---|---|---|
| Bulgarian | Qwen3.8-27B | 61.70 | 91.25 | not run | ready |
| Croatian | Qwen3.8-27B | 56.07 | 91.05 | not run | ready |
| Czech | Hy-MT2-7B | 57.99 | 92.51 | 91.87 | ready |
| Danish | Qwen3.8-27B | 64.81 | 90.83 | not run | ready |
| Dutch | Hy-MT2-7B | 55.42 | 88.63 | not run | ready |
| Estonian | Qwen3.8-27B | 52.73 | 90.29 | not run | ready |
| Finnish | Qwen3.8-27B | 52.98 | 91.67 | not run | ready |
| French | Hy-MT2-7B | 67.47 | 89.04 | 88.89 (23 Sep run) | ready |
| German | Hy-MT2-7B | 62.98 | 89.02 | 88.69 | ready |
| Greek | Qwen3.8-27B | 50.02 | 89.25 | not run | ready |
| Hungarian | Qwen3.8-27B | 52.39 | 89.37 | not run | ready |
| Italian | Hy-MT2-7B | 56.53 | 89.64 | not run | ready |
| Latvian | Qwen3.8-27B | 53.72 | 89.13 | not run | ready |
| Lithuanian | Qwen3.8-27B | 53.23 | 90.12 | not run | ready |
| Polish | Hy-MT2-7B | 50.10 | 90.54 | not run | ready |
| Portuguese | Hy-MT2-7B | 69.38 | 90.60 | 90.30 | ready |
| Romanian | Qwen3.8-27B | 61.81 | 91.05 | not run | ready |
| Slovak | Qwen3.8-27B | 55.25 | 90.54 | not run | ready |
| Slovenian | Qwen3.8-27B | 53.35 | 89.58 | not run | ready |
| Spanish | Hy-MT2-7B | 55.19 | 87.59 | 87.24 (23 Sep run) | ready |
| Swedish | Qwen3.8-27B | 64.65 | 91.07 | not run | ready |
| Irish | Qwen3.8-27B | 44.60 | 73.41 | not run | draft, needs a reviewer |
| Maltese | Qwen3.8-27B | 55.07 | 68.95 | not run | draft, needs a reviewer |
The bar: COMET-22 at least 85 and chrF++ at least 45, and COMET able to judge the language. 85 sits 2.6 points under the weakest language we already offered (Spanish, 87.59); no measured language falls between 74 and 89, so where exactly the bar sits between those does not change any route. COMET's encoder (XLM-R) was not pretrained on Maltese, so a Maltese COMET score cannot vouch for it and Maltese stays draft whatever its score. FLORES is Wikipedia-style prose; it says nothing about a language's legal or medical register, and no native speaker has reviewed any language.
The number lock in the added languages: the same 90 synthetic sentences translated by the routed model (Qwen3.8-27B) into Finnish, Hungarian, Greek, Bulgarian and Romanian, one error planted per variant (a digit, x10, a thousands separator, a unit, a dropped number, a negation, a date, a code). The word tables were written for this and fixed on the 30-sentence dev split; the 60-sentence test split was run once with the checker frozen:
| Language | Planted errors caught |
|---|---|
| Finnish | 210 / 214 |
| Hungarian | 212 / 215 |
| Greek | 214 / 216 |
| Bulgarian | 210 / 212 |
| Romanian | 218 / 221 |
| All five | 1,064 / 1,078 (98.7%) |
On the 300 clean test translations the frozen checker flagged 20: one real (a Hungarian "under 36 months" translated as "under 3 months") and 19 false (units the tables lacked: inches, the Greek "γρ.", the Bulgarian "г." read as a year; a "not" carried by a ne- prefix such as "Nepotrivit"). After fixes made from those (so no longer held out): 1,068 / 1,078 caught and 9 flagged, the same one real. Most misses are plants that did not change the meaning (removing the Greek "δεν" from "Κανένας ... δεν πέθανε" still leaves "no one") and the known case of a unit the translator added where the English had none.
Specialist candidates, measured the same way on a Mac (MLX, bf16, greedy, each model's own prompt) and not deployed: EuroLLM-9B-Instruct-2512 (Apache-2.0, 9.2B, all 24 EU languages) scores higher than Qwen3.8-27B in 14 of the 15 (Croatian lower), by -0.5 to +1.8 COMET in the ready languages, and by far in Irish (+7.9 COMET, +9.5 chrF++) and Maltese (+9.4 chrF++). Irish still falls under the bar with it. It is also about level with Hy-MT2-7B on Hy-MT2's own eight languages (-0.5 to +0.2 COMET), so one EuroLLM-9B in FP8 could serve all 24 in the slot Hy-MT2-7B holds today. Beside Hy-MT2 it would need about 12 GB more of GPU0, and the language pack is at its 30 GB ceiling, so it is proposed, not served. EuroLLM-22B (screened) is level with the 9B at 2.5x the memory; Salamandra-7B (screened) is under Qwen in 13 of 15. Hy-MT2-30B-A3B covers the same 33 languages as the 7B (no new EU language). NLLB-200 is CC-BY-NC and out.
| Language | Route: COMET / chrF++ | EuroLLM-9B: COMET / chrF++ | COMET gain |
|---|---|---|---|
| Bulgarian | 91.25 / 61.70 | 91.70 / 64.42 | +0.45 |
| Croatian | 91.05 / 56.07 | 90.53 / 54.56 | -0.52 |
| Czech | 92.51 / 57.99 | 92.07 / 56.15 | -0.44 |
| Danish | 90.83 / 64.81 | 91.63 / 67.80 | +0.80 |
| Dutch | 88.63 / 55.42 | 88.84 / 56.13 | +0.21 |
| Estonian | 90.29 / 52.73 | 92.07 / 57.10 | +1.78 |
| Finnish | 91.67 / 52.98 | 92.74 / 55.00 | +1.07 |
| French | 89.04 / 67.47 | 89.00 / 69.55 | -0.04 |
| German | 89.02 / 62.98 | 88.82 / 64.37 | -0.20 |
| Greek | 89.25 / 50.02 | 90.03 / 51.63 | +0.78 |
| Hungarian | 89.37 / 52.39 | 90.21 / 54.02 | +0.84 |
| Irish | 73.41 / 44.60 | 81.26 / 54.06 | +7.85 |
| Italian | 89.64 / 56.53 | 89.41 / 58.54 | -0.23 |
| Latvian | 89.13 / 53.72 | 90.96 / 57.65 | +1.83 |
| Lithuanian | 90.12 / 53.23 | 91.13 / 55.24 | +1.01 |
| Maltese | 68.95 / 55.07 | 71.66 / 64.45 | +2.71 |
| Polish | 90.54 / 50.10 | 90.49 / 50.56 | -0.05 |
| Portuguese | 90.60 / 69.38 | 90.09 / 68.86 | -0.51 |
| Romanian | 91.05 / 61.81 | 91.56 / 64.26 | +0.51 |
| Slovak | 90.54 / 55.25 | 91.30 / 58.12 | +0.76 |
| Slovenian | 89.58 / 53.35 | 90.45 / 55.70 | +0.87 |
| Spanish | 87.59 / 55.19 | 87.34 / 55.13 | -0.25 |
| Swedish | 91.07 / 64.65 | 91.66 / 67.30 | +0.59 |
Translation quality: the first six languages in detail
FLORES-200 devtest (CC-BY-SA 4.0), all 1,012 sentences, English into each language, the block's own prompt, greedy. chrF++ and BLEU from sacrebleu 2.6.0; COMET is Unbabel/wmt22-comet-da (Apache-2.0), reference-based, x100. Latency is per sentence at 12 requests in flight on a GPU shared with other evals.
| Language | chrF++ | BLEU | COMET-22 | p50 ms |
|---|---|---|---|---|
| German | 62.98 | 38.23 | 89.02 | 1,516 |
| French | 67.47 | 46.56 | 89.04 | 1,500 |
| Spanish | 55.19 | 29.64 | 87.59 | 1,825 |
| Italian | 56.53 | 29.56 | 89.64 | 2,400 |
| Dutch | 55.42 | 27.33 | 88.63 | 2,969 |
| Polish | 50.10 | 23.05 | 90.54 | 4,021 |
Same harness with Qwen3.8-27B (the 27B model we already serve), 23 Sep: German 62.80 / 88.61, French 68.77 / 88.89, Spanish 55.32 / 87.24 (chrF++ / COMET). Hy-MT2-7B is at that level with a quarter of the memory; Italian, Dutch and Polish were not run on Qwen3.8.
Planted drift: what the check catches
90 synthetic source sentences (labels, dosing, lay summaries, specifications; 30 dev, 60 test) translated by Hy-MT2-7B into six languages; one error planted per variant with regexes independent of the checker. Test split, run once with the checker frozen:
| Error planted | Caught |
|---|---|
| A digit changed (500 -> 600) | 318 / 318 |
| A number multiplied by ten | 292 / 292 |
| A thousands group written the English way (1.500 -> 1,500) | 10 / 10 |
| A unit swapped (mg/g, °C/°F, hours/days...) | 218 / 226 |
| A number and its unit dropped | 314 / 318 |
| The negation removed | 114 / 116 |
| A date changed | 36 / 36 |
| A code changed (model, EAN, phone) | 12 / 12 |
| All, six languages | 1,314 / 1,328 (98.9%) |
The same 360 clean translations: 9 flagged, of which 2 were real (a garbled French warning, a Spanish '2,000 W' left in English format) and 7 false (unit words the tables lacked, 'Betriebsstunden', 'avoid mixing' for 'do not mix'). After fixes made from those: 1,320 / 1,328 caught, 3 clean translations flagged; the 8 misses are units the translator added where the English had none.
The same check on professional translations (FLORES-200 devtest, 1,012 sentences, 250 with numbers): a flag is either a false alarm or a liberty the translator took (miles converted to km, '2,220' written as '2.200', a 'not' paraphrased away). Sentences flagged, frozen checker -> final:
| Language | Human translations flagged |
|---|---|
| German | 49 -> 46 of 1,012 |
| French | 48 -> 47 of 1,012 |
| Spanish | 63 -> 60 of 1,012 |
| Italian | 56 -> 55 of 1,012 |
| Dutch | 68 -> 65 of 1,012 |
| Polish | 66 -> 58 of 1,012 |
On real lay summaries
The English final drafts of the 12 held-out and 5 fresh trials of the lay summary eval (real ClinicalTrials.gov results) translated into six languages, each number traced to the results cells in that language. Numbers per language, 12 test trials:
| Language | Numbers traced (frozen) | After fixes | Planted errors caught |
|---|---|---|---|
| German | 658 / 684 (96.2%) | 662 / 672 (98.5%) | 410 / 410 |
| French | 665 / 704 (94.5%) | 671 / 679 (98.8%) | 395 / 396 |
| Spanish | 695 / 750 (92.7%) | 676 / 688 (98.3%) | 389 / 391 |
| Italian | 659 / 696 (94.7%) | 668 / 678 (98.5%) | 403 / 404 |
| Dutch | 661 / 679 (97.3%) | 668 / 679 (98.4%) | 420 / 421 |
| Polish | 664 / 692 (96.0%) | 669 / 681 (98.2%) | 409 / 410 |
Real errors the check found in the model's translations: 6 Polish sentences that dropped '0 out of' (reading 'in part 1 of 89 people died'), and one Spanish '30,000' in English format. Most untraced numbers in the frozen run were the checker misreading other languages ('249 des 325', group names, 'moins de 1 %'); the fixes are in the eval doc, and the 'after fixes' column is on the same trials.
Speech
Qwen3-ASR-1.7B on the first 100 FLEURS test sentences per language (CC-BY-4.0), language given, lower case and punctuation removed; digits are left as written, so a number spoken as a word counts as an error. TTS: 40 FLORES sentences without digits per language spoken by the house voice Ines (VoxCPM2, reference-clip mode) and transcribed by Qwen3-ASR; the WER measures both models together.
| Language | ASR WER (FLEURS, 100 sentences) | CER |
|---|---|---|
| German | 15.2% | 4.97% |
| French | 4.75% | 1.9% |
| Spanish | 2.36% | 0.91% |
| Italian | 8.6% | 2.93% |
| Dutch | 6.78% | 2.21% |
| Polish | 14.01% | 4.82% |
| Language | TTS round-trip WER | With check-and-retry |
|---|---|---|
| German | 23.35% | 14.19% (14 of 40 re-rendered) |
| French | 15.85% | 10.13% (10 of 40 re-rendered) |
| Spanish | 2.61% | 1.98% (2 of 40 re-rendered) |
| Italian | 20.15% | 6.75% (11 of 40 re-rendered) |
| Dutch | 6.73% | 5.38% (7 of 40 re-rendered) |
| Polish | 5.14% | 4.77% (4 of 40 re-rendered) |
The reference-clip mode keeps one voice across languages but sometimes garbles a line in German and French; the block transcribes every render and re-renders once when the transcript is more than 15% off the script. No native speaker has listened yet, so no language is advertised for speech. On sentences with numbers the picture is worse: rendering 20 dosing and label lines per language, the spoken check flagged 6 to 11 of 20. The renders garble symbols ("5 V, 2 A", "78 dB(A)", "31/12/2026") and sometimes invent content (a French "after 200 hours of use" came out as "every 15 to 20 years depending on the filter type"). Do not use this speech for safety text with numbers without the spoken check and a person listening.
Voices: two house voices (Ines, Tomas), each designed once from a written description with VoxCPM2's voice design; no recording of a person is used and uploaded voices are never cloned here. Each has a consent-ledger entry (use case 47) for narration, accessibility and audiobook use in the project decosa-language-pack; revoking it stops the voice.
Numbers spelled out before speech
Before speech, numbers, units, dates, times, percentages, amounts and codes are now spelled out in the language's own words by hand-written, deterministic tables ("5 V, 2 A" -> "fünf Volt, zwei Ampere", "31/12/2026" -> "trzydziestego pierwszego grudnia dwa tysiące dwudziestego szóstego roku"), with a map from each spoken span to its source span. The transcript of the render is read back into digits and compared with the source text, so a figure the voice invented is caught. VoxCPM2's own normalize option was tried first: it is English and Chinese only and writes "five V" into a German line, so it is not used. Same 20 dosing and label lines per language as before, plus 20 held-out lines never looked at while building; one render each, then the block's retry (one more take when the check fails). Real = the transcript lost or changed a number, date or negation, judged by reading (no listening).
| Condition (240 renders: 40 lines x 6 languages) | Flagged | Real errors | With an invented or changed figure |
|---|---|---|---|
| Before: text as written | 94 (39%) | 88 (37%) | 52 (22%) |
| VoxCPM2 normalize (20 lines only: 120 renders) | 46 of 120 (38%) | not labelled | 18 of 120 (15%) |
| After: spelled out, one take | 46 (19%) | 44 (18%) | 23 (10%) |
| After: spelled out, with the retry | 33 (14%) | 31 (13%) | 12 (5%) |
| Language | Real errors before (of 40) | After, with retry (of 40) | Verdict for spoken safety text |
|---|---|---|---|
| German | 12 | 3 | Not without a person listening |
| French | 20 | 11 | No: the voice still garbles a quarter of lines |
| Spanish | 16 | 4 | Not without a person listening |
| Italian | 15 | 5 | Not without a person listening |
| Dutch | 12 | 6 | Not without a person listening |
| Polish | 13 | 2 | Not without a person listening |
Held-out lines alone: 44 of 120 flagged before, 11 of 120 after with the retry. What remains is the voice itself: in reference-clip mode VoxCPM2 sometimes speaks a different sentence (a French label came out as "Ah non, pas les talibans..."), and codes such as "EN 71-3:2019" are still read wrong in most languages. The check flags these, and how many bad renders it misses is not measured (that needs people listening). No language is advertised for spoken safety text.
Meaning check
For what the number lock cannot see. Hy-MT2 translates each translated line back into English; Qwen3.8-27B through the gateway compares the English source with the translation and the back-translation and answers ok, check or error, with typed issues (negation, dropped clause, added fact, reversed roles, wrong entity, changed hedge) and the exact words in both texts. Two receipted model calls per line. Eval: 48 synthetic sentences (labels, leaflets, lay summaries, notices), one minimal edit per error type in the English, the edited English translated so the wrong translation is fluent, paired with the original source; 16 sources for tuning the prompt, 32 held out and run once.
| Planted error (held out, 6 languages) | Planted | Caught (check or error) | As error |
|---|---|---|---|
| Negation flipped | 106 | 102 | 100 |
| Clause dropped | 144 | 144 | 136 |
| Roles swapped | 90 | 89 | 85 |
| Entity changed | 132 | 132 | 130 |
| Hedge changed | 90 | 89 | 71 |
| All | 562 | 556 (98.9%) | 522 (92.9%) |
| Language | Caught / planted | Clean translations flagged (of 32) |
|---|---|---|
| German | 93 / 94 | 10 |
| French | 91 / 93 | 9 |
| Spanish | 94 / 95 | 7 |
| Italian | 92 / 93 | 10 |
| Dutch | 94 / 94 | 8 |
| Polish | 92 / 93 | 8 |
The six misses were edits the translator had silently undone, not missed errors. On the clean translations it flags about one line in five for a look (52 of 192; 4 as error): 7 of those were real faults in the model's own translation, 40 were false ("can" for "may", near-synonyms). On real lay summaries (5 ClinicalTrials.gov trials) it caught all 6 known Polish sentences that dropped "0 out of", and in a 240-sentence sample it found errors no number check sees: the inactive "Vehicle Cream" translated as a cream for cars in French, Spanish and Polish, "drop seizures" as a different seizure type in Italian, and "the trial" as "negotiations" in German and "court case" in Polish. 9 of its 14 ERRORs on that sample were false.
Per line: about 630 prompt and 50 completion tokens on the gateway (about $0.00026 at list price) plus one back-translation on our GPU; 7 s median, 10 s p95 with 8 lines in flight on a busy gateway. On in the GPSR pack; opt-in ("meaning": true) for lay summary Member-State versions, where 60 sentences x 6 languages take about 5 minutes more.
Grammar, gender and articles are out of scope ("den Ladegerät" passes). A model judgment, not a proof: a clean result does not mean the translation is right. Licence-clean QE models were checked and not used: CometKiwi and XCOMET are CC BY-NC-SA (non-commercial).
Our own QE model
A fast screen before the meaning check, and a model you can use commercially: the strong open QE models (CometKiwi, XCOMET) are non-commercial, so we trained our own. XLM-RoBERTa-large (MIT) reads the English line and its translation and returns the chance that the meaning changed, the likeliest kind of error and the words involved, in about 0.1 s on a CPU. Trained only on data that allows commercial use: Tatoeba, EMA and EU (DGT-TM) sentence pairs and our own sentences, with 46,768 planted errors (negation, dropped clause, added words, swapped names, terms or roles, changed numbers and hedges) in 100,478 pairs. Tested on the meaning check's own held-out set:
| Held-out planted set (562 errors, 192 clean, 6 languages) | Caught | Clean lines flagged | Per line |
|---|---|---|---|
| LLM meaning check alone | 556 (98.9%) | 52 (27%) | 7 s, $0.00026 |
| QE model alone | 476 (84.7%) | 6 (3.1%) | 0.1 s on a CPU |
| QE first, uncertain lines to the LLM (14% of lines) | 530 (94.3%) | 24 (12.5%) | about 1/7 of the LLM cost |
| Other tests | QE model | LLM meaning check |
|---|---|---|
| Planted number and negation drift (1,328 errors, 360 clean) | 1,309 caught, 9 clean flagged | not run (the number lock catches 1,314) |
| Real lay-summary translations (15 errors in 250 lines) | 5 caught (all dropped "0 out of"), 21 correct lines flagged | 12 caught, 50 flagged |
| Professional translations (FLORES-200, 1,200 lines) | 108 flagged (9%) | not run |
With "qe": "cascade" on POST /lang/meaning, lines the QE model is sure about are passed or flagged by it and only the rest go to the LLM; "only" skips the LLM. Thresholds were set on the dev split and not changed for the test sets. The cascade trades some recall (94.3% against 98.9%) for 86% fewer LLM calls and half the false flags; for safety text where every miss matters, keep the LLM on every line.
It also marks the words: F1 0.86 on changed numbers and codes, 0.53 on the meaning set (hedges best, entities worst). By type it catches dropped clauses 143 of 144, hedges 83 of 90, negations 88 of 106, entities 103 of 132 and swapped roles 59 of 90. Each call carries a model-call receipt with the sha256 of the weights.
Licences, read 27 Sep 2026: Tatoeba CC-BY 2.0 FR; EMA documents reusable for commercial purposes with attribution; DGT-TM under Commission Decision 2011/833/EU; our sentences CC0; translations by Hy-MT2-7B (Apache-2.0). FLORES-200 (CC-BY-SA) was used only to test, never to train. The weights are planned for release under Apache-2.0; not published yet.
Where it fails: wrong-sense words (a clinical "trial" translated as a court case: 0 of 9 real ones caught; the term bank handles these), plausible swaps ("the seller" became "the buyer" in all six languages, missed), negations that move within a sentence, and written-out abbreviations such as "Cream BID" as "twice a day", which cause most of its false flags on real lay summaries. English source and six target languages only. No native speaker has reviewed the labels.
Term bank
The meaning check's real catches were domain terms: a placebo 'Vehicle Cream' translated as a cream for cars in French, Spanish and Polish, 'the trial' as negotiations in German and a court case in Polish. The term bank is terminology memory per customer and per domain: required renderings, forbidden ones (the wrong sense), a domain tag, a note and where each entry came from.
- Each customer (API key) keeps its own bank. Public packs: two hand-made ones (clinical-trial lay summaries, GPSR product safety, six languages, with forbidden wrong senses) and nine generated from EU legislation in up to 24 languages. A request glossary still wins over the bank.
- Required renderings go into Hy-MT2's own terminology prompt. After translation, code checks each term: a required form must be there and no forbidden one. A wrong-sense rendering is flagged with its span, and the line is translated once more with the rule named; if it still fails, it stays flagged.
- Every change to a bank is a signed, append-only event. The bank version (the customer's sequence number and head hash, and each pack's hash) goes into every model call's receipt and the signed translation receipt, so a customer can show which version a translation used.
- When the meaning check or a reviewer catches a wrong term, the catch becomes a proposal. Proposals never apply until a person on the customer's side accepts them.
- On by default for lay summary Member-State versions (clinical-trials pack) and the GPSR pack (product-safety pack); on request for /lang/translate. A run can add packs: a GPSR pack for a device adds the MDR or IVDR definitions with term_bank: {"add": ["eu-mdr"]}. A term with no entry in a language is not enforced there, and the run's coverage says which.
The five fresh trials of the lay summary eval, translated again without a bank, with the clinical-trials pack, and with the pack plus one accepted customer entry ('drop seizure'). The errors the meaning check had found:
| Error | No bank | Pack | Pack + one accepted entry |
|---|---|---|---|
| French 'Vehicle Cream' as 'Crème pour véhicules' | reproduced | 'crème placebo' | same |
| Spanish 'Vehicle Cream' as 'crema para vehículos' | reproduced | 'crema con placebo' | same |
| Polish 'Vehicle Cream' as 'kremem do pojazdów' | reproduced | 'kremem placebo' | same |
| German 'The trial' as 'Die Verhandlungen' | reproduced in 1 of 2 runs | 'Die Studie' | same |
| Polish 'The trial' as 'Proces' | reproduced | 'Badanie' | same |
| Italian 'drop seizures' as 'crisi ipertoniche' | reproduced | 'crisi' (wrong type gone, 'drop' lost) | 'crisi con caduta' |
5 of the 6 term errors prevented by the public pack alone and the sixth half-fixed; 6 of 6 with one reviewed customer entry. A seventh error (a garbled '2; 72 mesi') is not a terminology error and did not recur. The first run also showed two false matches (a name, 'EU Clinical Trials website', and the compound 'anti-seizure'), fixed after that run. On the second run the bank made no sentence fail and added 8 retries in about 2,330 calls; 27 more sentences were marked 'check' because Hy-MT2 then also wrote 'BID' out as 'twice a day', which the number lock flags although it is right. On 249 sampled sentences the meaning check flagged 58 without the bank and 61 with it; reading the 11 that got worse: 6 are the check objecting to 'placebo' for 'vehicle' and similar pack terms, 2 are real new errors, 3 are noise.
60 synthetic sentences (clinical-trial and product-safety lines with terms such as trial, vehicle, seizure, recall, withdrawal, sponsor), 20 dev and 40 test, in six languages. The packs were frozen after the dev run. Test, run once (240 translations, 435 term occurrences):
| No bank | With the pack | |
|---|---|---|
| Terms rendered as the pack requires | 315 (72.4%) | 434 (99.8%) |
| Wrong sense (court case, car, recall and withdrawal swapped), by reading | 15 | 0 |
| Meaning shift across a defined distinction (adverse event as side effect, recall as withdrawal), by reading | 14 | 0 |
| Non-standard or wrong term a reader can decode, by reading | 17 | 0 |
| Model calls | 240 | 241 |
| Grammar or spelling broken by a forced term / content added | - | 8 / 1 |
The 'with the pack' column is graded by the rules that steered it, so every changed term was read, by the same agent that wrote the packs and the sentences. The packs enforce the regulation's defined terms, so 67 acceptable everyday synonyms ('patrocinador', 'Ethikkomitee') were replaced too, and 6 correct translations were rejected by forms that were too narrow. The meaning check flagged 79 translations without the bank and 92 with it: of the 48 that got worse, 30 are the check objecting to the pack's own terms, 2 are real new errors, 16 are noise.
Every EU regulation defines its terms in one article, numbered the same in all 24 language versions. A script reads the Official Journal's own XML (Formex, from the Publications Office by CELEX number) for each language, takes the defined term at each point, and lines the versions up by point. A language whose points do not line up (a different count, labels, or the numbers each definition cites) is left out rather than shifted. The same files give the same pack, byte for byte.
| Public pack | Terms | Built from | Licence |
|---|---|---|---|
| Clinical-trial lay summaries | 12 · 6 languages | EU Clinical Trials Regulation 536/2014 Art. 2 in six languages; IATE entries (placebo, vehicle, side effect, adverse event, seizure) and their wrong senses; three catches of the meaning check | Reuse with acknowledgement (Commission Decision 2011/833/EU) |
| GPSR product safety | 16 · 6 languages | General Product Safety Regulation 2023/988 Art. 3 in six languages; IATE for warning, instructions for use, batch number | Reuse with acknowledgement (Commission Decision 2011/833/EU) |
| Clinical Trials Regulation definitions (eu-ctr) | 35 · 24 languages | Regulation (EU) No 536/2014, Art. 2 (CELEX 32014R0536) | Reuse with acknowledgement (Decision 2011/833/EU, EUR-Lex legal notice) |
| GPSR definitions (eu-gpsr) | 28 · 24 languages | Regulation (EU) 2023/988, Art. 3 (CELEX 32023R0988) | same |
| Medical Devices Regulation definitions (eu-mdr) | 73 · 24 languages | Regulation (EU) 2017/745, Art. 2 (CELEX 32017R0745) | same |
| IVDR definitions (eu-ivdr) | 76 · 24 languages | Regulation (EU) 2017/746, Art. 2 (CELEX 32017R0746) | same |
| AI Act definitions (eu-ai-act) | 68 · 24 languages | Regulation (EU) 2024/1689, Art. 3 (CELEX 32024R1689) | same |
| Digital Services Act definitions (eu-dsa) | 24 · 24 languages | Regulation (EU) 2022/2065, Art. 3 (CELEX 32022R2065) | same |
| Digital Markets Act definitions (eu-dma) | 33 · 24 languages | Regulation (EU) 2022/1925, Art. 2 (CELEX 32022R1925) | same |
| GDPR definitions (eu-gdpr) | 27 · 24 languages | Regulation (EU) 2016/679, Art. 4 (CELEX 32016R0679) | same |
| ESRS glossary (eu-esrs) | 199 · 22 languages (Italian and Lithuanian sort the glossary differently and are left out) | Delegated Regulation (EU) 2023/2772, Annex II Table 2 (CELEX 32023R2772) | same |
563 terms and 12,497 renderings. Checked by reading 428 of them (8 terms per pack, 6 languages each, every language covered) against the English: 0 wrong terms; the first build had 3 Hungarian terms with a stray article, fixed before that final reading. A reliable IATE term agrees with the pack for 78.6% of the renderings IATE has at all (4,302 of 5,472); most of the rest are other senses or older wording. Finnish and Czech name terms in an oblique case, so their prompt form comes from IATE where one fits. CSRD was not built: its definitions are amendments to the Accounting Directive; the ESRS glossary carries the reporting terms. No linguist has reviewed the generated packs.
IATE, the EU institutions' terminology database, is kept offline: the official full export (the download is from February 2019, 934,921 entries), filtered to health, law, trade, product safety, finance and information-technology domains: 351,505 entries in 24 languages. It only suggests: a proposal whose right rendering is unknown shows IATE's reliable terms (reliability 3-4, never deprecated ones) to the reviewer, and GET /lang/termbank/iate answers a reviewer's lookup. Nothing from IATE is enforced until a person accepts it.
The meaning check now knows the bank: the terms whose required rendering is in the translation are listed to the judge as correct by definition, and an issue that is only about such a term is set aside (and kept in the output). Only sentences that contain a bank term were judged again, twice on the same day, without and with the list:
| Sentences with a bank term | Flagged before | Flagged after | Flagged only for a bank term, before -> after |
|---|---|---|---|
| Planted test, translated with the pack (238) | 90 | 57 | 51 -> 0 |
| Planted test, translated without a bank (199) | 59 | 66 | 16 -> 0 |
| Real lay summaries, pack (96) | 30 | 19 | 13 -> 0 |
| Real lay summaries, pack + one accepted entry (96) | 28 | 28 | 11 -> 0 |
| Planted meaning errors that contain a bank term (54) | 54 caught | 54 caught | - |
Every known real error is still flagged: the car cream left in a translation made without the bank, 'crisi ipertoniche' for drop seizures, 'Circa 150', the German dose word that loses 'twice', 'chauffeurs' for heaters and the added Dutch clause. The judge is told to check every other word as strictly as before, so on translations made without the bank it flags a little more (59 -> 66): some are real wrong terms outside the bank (Spanish 'retirada' for a recall), most are the usual noise. Tuning was on a separate dev set, plus one fix after the test run, disclosed: the first rule also set aside 'serious' becoming 'mild' next to a term, and now only an article, or a word the back-translation kept, may sit next to it.
Not reviewed by a linguist or native speaker. Generated packs enforce each regulation's defined words wherever they occur once a pack is chosen, including generic ones such as product, risk and consumer, so choose the packs that fit the document. Not covered yet: 'treatment arm' (no source has the trial sense), 'drop seizure' (a customer entry in the demo bank). MedDRA (licensed) and CDISC terminology are not used. IATE's download is from 2019, so newer terms (AI Act, DSA, GPSR) are not in it.
Browse the bank
The public packs and the read-only demo bank, live from the API. With an API key you keep your own bank: add entries, review proposals, pin a version.
Loading…
API
| POST /lang/translate | {text, source_lang?, targets: [up to 23 of the 24 EU languages], glossary?, keep?, term_bank? (packs, add, tenant, at), stream?} -> per language the text, the model that served it and its tier (ready, or draft needing a reviewer), per line the flags with source spans, receipts and a signed translation receipt. Any token or dk_ key. |
| POST /lang/check | {source, target, target_lang, glossary?, keep?} -> the same checks on a translation you bring. No model, no token. |
| POST /lang/meaning | {source (English), target, target_lang} -> per line ok / check / error with typed issues, spans, a reason and the back-translation. Two receipted model calls per line. Any token or dk_ key. Add "qe": "cascade" to let the QE model decide the lines it is sure about, or "only" for no LLM call. |
| POST /lang/speak | {text, language, voice: a house voice, purpose?, normalize?: true} -> WAV in a consented house voice; numbers, units, dates and codes spelled out first (the spoken text and a span map come back), the spoken check reads the transcript back into digits against the source; consent decision and receipts. |
| GET /lang/termbank/packs, /lang/termbank/demo | The public term packs (licences, sources, the source act's CELEX and ELI, languages, entries) with each use case's default and optional packs, and the read-only demo bank. No token. |
| GET /lang/termbank, /lang/termbank/log; POST /lang/termbank/verify | Your own bank (a dk_ key): entries, pending proposals, version {seq, head}; its signed event log; a check of the log and of the bank version a translation receipt names. |
| POST /lang/termbank/entries, /proposals, /proposals/{id}/accept|reject | Add, update or retire an entry; queue a proposal from a meaning-check result or a reviewer's fix; a person accepts or rejects it. Every change is a signed, append-only event. |
| GET /lang/termbank/iate?term=&langs= | Reliable IATE renderings of an English term per language, from the offline index, for a reviewer. Suggestions only; any token. |
| POST /lang/transcribe | {audio_b64, language?, context?} -> text with a receipt. |
| POST /lang/verify | {receipt, source?, texts?} -> signature and text checks of a translation receipt. |
| GET /lang/info, /lang/voices, /lang/samples | Languages and the route per language (model, tier, FLORES scores), models with revisions and weight hashes, checks, limits; the house voices and their consent status. |
| POST /lang/qe | {source (English), target, target_lang: de|fr|es|it|nl|pl} -> per line the chance of a meaning error, ok / check / error, the likeliest error type and the words, from our own QE model on a CPU. One model-call receipt; no generated tokens charged. |
GPU memory and licences
| Part | Measured |
|---|---|
| Hy-MT2-7B, vLLM, gpu-memory-utilization 0.18 (weights 14.0 GiB + KV) | 18.3 GB |
| Qwen3-ASR-1.7B + VoxCPM2, both loaded (each unloads after 15 min idle) | 10.4-11.5 GB |
| Qwen3.8-27B for the fifteen added languages | 0 GB new: the model already served on the other card, through the gateway |
| All three at once | 29.7 GB (ceiling 30 GB) |
| EuroLLM-9B-Instruct-2512, proposed (not deployed) | about 12 GB in FP8 beside Hy-MT2 (estimate, not measured on the card), or in Hy-MT2-7B's own 18 GB slot as a replacement: no room under the ceiling otherwise |
| QE model (CPU only, fp32) | 0 GB GPU · 2.9 GB RAM |
| Model | Licence and revision |
|---|---|
| Hy-MT2-7B (tencent/Hy-MT2-7B) | Apache-2.0, no territory clause (LICENSE.txt read 27 Sep 2026) · rev 9b0eb4e8 |
| Qwen3.8-27B (Qwen/Qwen3.8-27B), the fifteen added languages | Apache-2.0 (model card read 27 Sep 2026) · rev 1d4bf0f2 · served through the gateway |
| Qwen3-ASR-1.7B (Qwen/Qwen3-ASR-1.7B) | Apache-2.0 (model card) · rev 7278e1e7 |
| VoxCPM2 (openbmb/VoxCPM2) | Apache-2.0 (model card; GitHub repo Apache-2.0) · rev 32279eff |
| COMET scorer (Unbabel/wmt22-comet-da) | Apache-2.0 · eval only |
| Candidates measured, not served: EuroLLM-9B/22B-Instruct-2512, Salamandra-7B-instruct, MADLAD-400-3B-MT | Apache-2.0 each (model cards read 27 Sep 2026). Excluded: NLLB-200 (CC-BY-NC-4.0), SalamandraTA-7B-instruct (GPL-3.0) |
| QE model (decosa-qe-xlmr-large v2, ours) | base XLM-RoBERTa-large MIT; trained on Tatoeba (CC-BY 2.0 FR), EMEA (EMA reuse notice), DGT-TM (Decision 2011/833/EU), our CC0 sentences; weights: Apache-2.0 planned, unreleased |
Non-LLM model calls are attested by decosa-api's own key today (an operator attestation, not a proof); the statement is our gateway's model-call attestation format, so the gateway can countersign it once its model-call route exists.
Built on it
GPSR listing pack · built
Warnings and safety information in up to five of the 23 EU languages for an EU listing, with numbers and names checked.
EU trial lay summary: Member-State versions · built
Each sentence translated into any EU language and its numbers traced to the results cells again in that language.
Author-voice audiobook in translation · not built
A sample chapter in a house voice (the consent gate allows audiobook use); an author's own voice needs its own consent entry.
Consented dubbing beyond Spanish, language-access QA, streamer captions · candidates
The same Translator, checker and ASR.
Not done
- No native speaker has read a translation or listened to a voice; no language is advertised for speech until one has.
- Speech is not ready for unattended safety text: with numbers spelled out and one retry, 5-28% of dosing and label renders per language still lose or change a number, date or negation (French worst). The spoken check flags them; nobody has measured what it misses by listening.
- The meaning check is a model judgment: it flags about one correct line in five for a look and does not check grammar, gender or articles.
- The number check in the fifteen added languages was tuned on a 30-sentence dev split and tested on five of them (Finnish, Hungarian, Greek, Bulgarian, Romanian); the other ten have word tables but no planted-error run. The meaning check was measured on the first six languages only.
- Negations are not read for Czech, Slovak, Lithuanian, Latvian (a ne- prefix on the verb) and Maltese (a -x suffix); a translation that loses a 'do not' in those languages is not flagged.
- Irish and Maltese are drafts: Qwen3.8-27B scores COMET 73.4 in Irish and 69.0 in Maltese (and COMET cannot judge Maltese). EuroLLM-9B does clearly better on both but has no GPU room yet.
- Number words above 20 (other than tens and 100) are not read.
- The 30B-A3B translation model and an A/B against Qwen3.8-27B on the same lines were not run.
- Term bank: the public packs enforce the regulations' defined terms, including where an everyday synonym would do, and no linguist has reviewed them; forced terms sometimes break agreement (8 of 240 test translations).
- The QE model misses wrong-sense words and more role and entity swaps than the LLM check; its thresholds come from synthetic dev data and flag 9% of correct real lay-summary lines.