{"schema_version":"1","site":"https://decosa.ai","page":"/metrics","how_we_measure":[{"id":"evals","heading":"Quality evals","text":"Each tool has an eval of its own: cases with planted problems or known answers, scored by a script or a judge model. Where there is a test split, the prompts were tuned on the dev split only and the test split was run once the code was frozen. We show the test number, the dev number where it matters, and the error or false-positive rate next to the hit rate."},{"id":"held-out","heading":"Held out, or not","text":"“Held out” means the scored cases were never used to change prompts or code: a test split kept apart, or a public benchmark we did not tune on. Many of our evals are not held out; the page says so. Those numbers show the pipeline works on the cases we wrote, not how well it generalises."},{"id":"synthetic","heading":"Synthetic data","text":"Most cases are synthetic: written for the eval, often by the same person who wrote the prompts, because real clinical notes, privileged documents or customer calls cannot be published. Synthetic cases tend to be cleaner and more blatant than real ones. Expect real-world numbers to be lower."},{"id":"nightly","heading":"Nightly smoke checks","text":"Every night a script calls each tool's hosted API with its own sample input through a dedicated key and checks the answer and its receipts. Pass or fail, latency, receipts, model calls and cost come from that run. Cost is estimated at the gateway list price. A smoke check proves the pipeline runs end to end; it is not a quality score."},{"id":"receipts","heading":"Receipts","text":"Model calls on the hosted route return signed receipts: which model ran, on which inputs (by hash), with what output. Receipts prove what ran, not that the answer is right."},{"id":"not-claimed","heading":"What we don't claim","text":"No eval here is an independent audit, a clinical or legal validation, or a certification. Sample sizes are small. Numbers from a proxy task (for example a clinical-notes benchmark shown for a sales tool) are labelled as proxies and are never the headline. Where nothing is measured yet, we say “not measured yet”."}],"nightly":{"live_url":"https://api.decosa.ai/verify/status","note":"Live nightly results (pass/fail, latency, receipts, model calls, tokens, cost, 30-day history) are served by the API, not baked into this file. `verification.hosted` below is the last manual QA sweep and is the fallback when the API does not answer."},"counts":{"use_cases":86,"with_eval_metrics":85,"held_out":69,"with_rehearsal_bundle":85},"use_cases":[{"id":"clinical","num":"01","name":"Visit copilot","status":"live","industries":["healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Medical-term miss rate, live ASR (Voxtral Mini 4B Realtime)","value":"8.4%","unit":null,"n":57,"split":"heldout","note":"146 of 1,741 lexicon terms; Nemotron-3.5 streaming 12.7%; WER 13.2 (Nemotron-3.5 streaming 11.8; MOSS-TD pass 2 10.3)"},{"name":"Speaker labels during the visit, word-level role accuracy (rolling MOSS windows)","value":"99.71%","unit":null,"n":11,"split":"heldout","note":"PriMock57, 110 min of audio; whole-recording pass 99.97%; worst consultation 98.1%"},{"name":"Considerations: warning-feature encounter recall (synthetic, held-out)","value":"14/15 and 13/15","unit":null,"n":15,"split":"heldout","note":"two runs; dev 7/7 both; topic recall held-out 13/16 and 12/16"},{"name":"Guidance: false-alarm encounters","value":"1/16 and 1/16","unit":null,"n":16,"split":"heldout","note":"both flags outside the guideline pack (tea-coloured urine on a statin; drowsy driving), shown as 'no bundled guideline'"},{"name":"Guidance: directive wording shown","value":"0","unit":null,"n":null,"split":"heldout","note":"all runs, the live lint"},{"name":"Guidance items judged useful at that moment (blind clinician judge)","value":"checklist 95/145 (66%); on real GP consultations 20/31 (65%)","unit":null,"n":425,"split":"heldout","note":"blind clinician judge: 76 moments in 38 synthetic visits and 24 in 12 PriMock57 consultations; warning-feature items 4/10 and 1/5 useful, the rest already covered; history elements 11% and 18% useful, so they fold away during the visit; 3 of 523 items judged harmful, all synthetic"},{"name":"Note sentences supported by the human transcript (blind judge, held-out)","value":"289/308 (93.8%)","unit":null,"n":308,"split":"heldout","note":"PriMock57, 11 new consultations; the 28 Sep writer 229/254 (90.2%) on the same visits; writer-caused misses 12 -> 3; 16 misheard by speech recognition, which a transcript self-check cannot see"},{"name":"WH-380-E boxes filled right (held-out FMLA visits written blind)","value":"91.4% (85/93)","unit":null,"n":8,"split":"heldout","note":"28 Sep code 79.6% on the same visits; 85.3% vs 75.7% by code scoring alone (free text judged blind otherwise); false fills 10 -> 2. Dev: work/school note 90%, instructions 97%"},{"name":"Signatures filled by a paperwork draft","value":"0","unit":null,"n":28,"split":"heldout","note":"signature and signing-date boxes"},{"name":"Guidance checklist items the GP went on to ask (PriMock57, real GPs)","value":"25/43 (58%)","unit":null,"n":43,"split":"heldout","note":"history elements 88/107 (82%); what the lane listed that the clinician asked later in the same visit"}],"dataset":"PriMock57 (57 recorded mock primary-care consultations, CC BY 4.0) for speech, speaker labels and note grounding; 38 synthetic US visits with planted warning-feature labels (written and labelled by separate agents) for guidance; 12 synthetic visits with paperwork truth written blind by a separate agent. Judges: Claude Code Opus 5.5 as blind sub-agents.","held_out":true,"caveats":["Synthetic visits and mock consultations, not real clinic audio.","Two-speaker visits only.","The judges are models (Claude Code Opus as blind sub-agents), not clinicians.","The guidance was tested on a 26-topic pack.","Costs are at list price; eval runs used the direct route (same weights), the e2e runs the gateway."],"date":"2026-09-26","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card, live pass only","evidence":[{"metric":"ACI-Bench ROUGE-L (Qwen3.8-27B FP8, human transcript)","value":"34.2","source":"scribe-bench wiki models.md / RESULTS.md"},{"metric":"PriMock57 note composite, vanilla pipeline (streaming ASR -> Qwen3.8-27B, official weights, test 37)","value":"41.01","source":"scribe-bench wiki vanilla-vs-best"},{"metric":"Medical-term miss / WER, live ASR (Voxtral Mini 4B Realtime, PriMock57, 57 visits)","value":"8.4% / 13.2 (Nemotron-3.5 streaming: 12.7% / 11.8)","source":"scribe-bench asr_score on PriMock57 (57 visits) through the live realtime endpoint; eval results file asr-voxtral-primock57.json (2026-09-23)"}]},{"tier":"standard","label":"Standard · one 96 GB Blackwell card, two passes","evidence":[{"metric":"Medical-term miss / WER, pass 2 MOSS-TD vs the Voxtral live pass (PriMock57, 57 visits)","value":"8.4% / 10.3 vs 8.4% / 13.2","source":"scribe-bench RESULTS.md (MOSS-TD); eval results file asr-voxtral-primock57.json (Voxtral, 2026-09-23)"},{"metric":"DER / word speaker misattribution, MOSS-TD","value":"11.4 / 1.0%","source":"scribe-bench RESULTS.md, wiki decoder-finding"},{"metric":"Composite, two-pass + role map + Qwen3.8-27B minus vanilla (official weights, test 37)","value":"+2.0 [-1.0, +5.7], not resolved; misattributions -0.05","source":"scribe-bench wiki vanilla-vs-best (measured with Sortformer + Parakeet as pass 2)"},{"metric":"Verifier recall on injected errors / flags on clean, Qwen3.8-27B judge","value":"99.1% / 6.5%","source":"scribe-bench wiki verifier"}]},{"tier":"best","label":"Best · adds DeepSeek V4 Flash as the note writer on 2x 96 GB","evidence":[{"metric":"ACI-Bench base ROUGE-L / term precision / plan recall","value":"35.8 / 70.3 / 93 (Qwen3.8-27B ROUGE-L 34.2)","source":"scribe-bench wiki models.md"},{"metric":"PriMock57 cited note, official weights (test 37): grounded / term precision","value":"91.33% / 31.06 (vanilla 88.62% / 26.41)","source":"scribe-bench wiki vanilla-vs-best, citations-and-verifiability"},{"metric":"Composite vs vanilla, official weights","value":"+1.25 [-3.76, +6.23], not resolved; plan recall -8.1 [-13.8, -1.7]","source":"scribe-bench wiki vanilla-vs-best"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":60000,"p95_ms":71600,"runs":null,"receipts_per_run":80,"cost_per_run_usd":0.07},"selfhost":{"date":"2026-09-29","result":"pass","method":"fresh clone of the branch into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume and DECOSA_CLINIC_PROFILE, pointed at the model servers already running on our server (Qwen3.8-27B, Voxtral, MOSS diarizer, M17) instead of starting new ones; then torn down","notes":"The image builds and starts; /healthz ok with asr, llm and diarize true. The copilot sample at 2x passed end to end: 28 speaker turns, guidance, a self-checked note (28 sentences, M17 on), codes from the assessment, three paperwork drafts with the clinic profile from the environment, 0 signatures; done 44.5 s after the audio on a GPU shared with our evals. Model-server startup itself was not re-verified."},"known_limits":["Hosted is for synthetic visits only; real visits must be self-hosted (no BAA yet).","Speaker labels were tested on two-speaker visits; a third speaker is labelled Other.","The self-check compares the note with the visit's own transcript, so a word the recogniser misheard passes it. The note marks sentences that may rest on one (where the two recognisers disagree, or a word is unknown): 8 of 17 such sentences on held-out visits, with 8% of good sentences also marked. Check names, numbers and yes/no answers against the cited turn.","WH-380-E drafts still need every box checked (91% of filled boxes right on held-out visits; vague essential-function wording and incomplete date lists are the usual misses); third-party insurer FMLA forms are not bundled yet.","No EHR write-back yet: copy the note and the drafts, or use the API.","Guidance is a 26-topic public-domain pack plus the CMS history elements; a missing item does not mean nothing is missing."],"receipt_coverage":"full"},"cost_per_run_usd":0.07,"rehearsal_bundle":{"url":"/samples/clinical.zip","checks":21,"bytes":676298},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Pass 1: live streaming transcript for the in-visit view (no speakers)","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note cites","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mapping","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa-note-detail-checker-modernbert-large (M17, our own model)","role":"Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail error","license":"Apache-2.0","hf_repo":"decosaai/decosa-note-detail-checker-modernbert-large"}],"licence":"permissive","links":{"metrics":"/metrics/clinical","page":"/clinics/visit-copilot","json":"/use-cases/clinical.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"code","num":"03","name":"Private code assistant","status":"live","industries":["software"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Coding benchmark, 5 core tasks, direct API (Qwen3.8-27B, vLLM NVFP4 + MTP)","value":"478 / 500","unit":null,"n":5,"split":"test","note":"single run"},{"name":"Same weights on other engines","value":"vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (/500)","unit":null,"n":5,"split":"test","note":"engine bugs, not the model"},{"name":"Code smoke on the exact pinned stack","value":"12/12","unit":null,"n":12,"split":"test","note":null},{"name":"Best tier (DeepSeek V4-Flash), agent harness","value":"495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses) / 500","unit":null,"n":5,"split":"test","note":"measured on the DSpark serving build, not the NVFP4 kit this tier uses"}],"dataset":"coding-agent-bench (the owner's own coding benchmark, 5 core tasks scored /500), results/RESULTS.md on our server, plus the 12-task code smoke from the model benchmark page.","held_out":false,"caveats":["The owner's own benchmark with only 5 tasks; not an independent public benchmark.","Single runs; run-to-run variation not measured.","A 32 GB card (lite tier) was not measured specifically.","The best-tier figures were measured on a different serving build than the one this tier ships."],"date":"2026-09-23","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · runs on one 32 GB Blackwell card","evidence":[{"metric":"Coding benchmark, 5 core tasks, direct API (/500)","value":"478 (same weights; measured on a 96 GB card, single run)","source":"coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16"},{"metric":"On a 32 GB card specifically","value":"not measured yet","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one 96 GB card (hosted demo)","evidence":[{"metric":"Coding benchmark, 5 core tasks, direct API (/500)","value":"478 (vLLM NVFP4 + MTP, single run)","source":"coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16"},{"metric":"Same weights on other engines (/500)","value":"vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (engine bugs, not the model)","source":"coding-agent-bench results/RESULTS.md"},{"metric":"Code smoke on the exact pinned stack","value":"12/12","source":"Decosa model benchmarks (Sep 2026)"}]},{"tier":"best","label":"Best · two 96 GB cards","evidence":[{"metric":"Coding benchmark, 5 core tasks, agent harness (/500)","value":"495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses); single runs","source":"coding-agent-bench results/RESULTS.md, official_prime_results.json and ds4_*_results.json, 2026-08-15/16"},{"metric":"Same suite, direct API (/500)","value":"459 uncapped (single run)","source":"coding-agent-bench results/RESULTS.md, ds4_api_uncapped_results.json"},{"metric":"Caveat","value":"measured on the DSpark serving build, not on the NVFP4 kit this tier uses; NVFP4 kit not measured yet","source":"coding-agent-bench results/RESULTS.md"}]},{"tier":"wanted","label":"Wanted · the two most-used open coding models","evidence":[{"metric":"Coding benchmark, 4 tiebreaker tasks, Prime Agent harness (/400)","value":"372.3 three-run mean (389, 359, 369), community TR3 4bpw build, 8k thinking budget","source":"coding-agent-bench results/RESULTS.md, h2h/glm53_tr3_t8k_tb_r1-3.json, 2026-08-28"},{"metric":"DeepSeek-V4.1-Flash on the same benchmark","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":4268,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.00046},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers","notes":"Verified on 2026-09-25: the fallback llm image builds, the compose file validates, the api starts and the sample passes end to end (attested receipts, p50 1.5 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Tool calling was not re-verified: the local server ran without the tool-parser flags."},"known_limits":["Hosted answers stop at 2,048 generated tokens (finish_reason \"length\"); self-host for longer outputs.","Tool calling works on the hosted route (checked on production 2026-09-30: a request with `tools` returns `tool_calls` and a receipt, and the follow-up turn with the tool result is answered). Each call is still one request of at most 2,048 generated tokens.","The hosted route computes the whole answer before streaming it, so text arrives in one burst.","Hosted p50 is for an answer of about 340 tokens (7 calls on 2026-09-30, fastest 2.5 s, slowest 5.1 s); under load on 2026-09-25 the same call took up to 37 s.","Token counts on hosted receipts are the gateway's metering, which on 2026-09-25 overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"receipt_coverage":"full"},"cost_per_run_usd":0.00046,"rehearsal_bundle":{"url":"/samples/code.zip","checks":9,"bytes":1347},"models":[{"name":"Qwen3.8-27B (NVFP4)","role":"Coding model (chat, edit, agent tool calls)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/code","page":"/tools/developer/code","json":"/use-cases/code.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"studio","num":"04","name":"Decosa Studio","status":"live","industries":["creative-media","personal-family"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[],"dataset":"No quality eval. One reproducibility check on our server: same-seed re-render is bit-identical for MiniMax-Music3 and Qwen-Image-2512, not identical for ACE-Step.","held_out":false,"caveats":["Output quality not measured yet on any tier.","The same-seed check shows reproducibility, not quality."],"date":"2026-09-23","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · drafts on one 24–32 GB card","evidence":[{"metric":"output quality","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, about 50 GB of one 96 GB card","evidence":[{"metric":"output quality","value":"not measured yet","source":null},{"metric":"same-seed re-render","value":"bit-identical for MiniMax-Music3 and Qwen-Image-2512; not identical for ACE-Step","source":"measured on our server 2026-09-23"}]},{"tier":"best","label":"Best · self-host only, a whole 96 GB card","evidence":[{"metric":"output quality","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","evidence":[{"metric":"render time per clip against one card with offload","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"partial","p50_ms":null,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":null},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of decosa-api; the api image plus ffmpeg running the repo's own studio worker against an already-running ComfyUI on the same box (no render or ComfyUI containers built, no new model loads).","notes":"Verified on 2026-09-25: the image builds, the service starts, and the prompt's MiniMax-Music3 smoke job (seed 5501) finished in 40 s against a local ComfyUI equivalent to the documented one; model-server startup and the image and video paths were not re-verified. Two runs of the same seed gave bit-identical decoded PCM. An api restart mid-render marked that job failed with a reason and ran the queued one (72 s). Fixed on the way: the api image left out services/ (every render failed), jobs were lost on restart."},"known_limits":["Renders share one GPU with the live audio demos: a job pauses while a live session runs.","Each token or key can queue 3 jobs; hosted video needs an API key and fal credit.","Receipts for renders are our signed statement of model, seed, weights and output hash; nobody re-renders to check them.","Wan2.1 video takes about 35 minutes per 5 s clip on the shared GPU."],"receipt_coverage":"partial"},"cost_per_run_usd":null,"rehearsal_bundle":{"url":"/samples/studio.zip","checks":11,"bytes":964918},"models":[{"name":"MiniMax-Music3","role":"Music generation (songs with vocals and lyrics)","license":"MiniMax-Music3 Community License","hf_repo":"MiniMaxAI/MiniMax-Music3"},{"name":"Qwen-Image-2512","role":"Image generation","license":"Apache-2.0","hf_repo":"Qwen/Qwen-Image-2512"},{"name":"MiniMax H3 Max (via fal)","role":"Hosted video (default): 5 s clips with audio","license":"H3's own licence excludes the US; used here only through fal's hosted endpoints, listed as commercial use under fal's MiniMax partnership","hf_repo":null},{"name":"Kokoro-82M","role":"Text to speech (preset voices)","license":"Apache-2.0","hf_repo":"hexgrad/Kokoro-82M"},{"name":"IndexTTS-2.5","role":"Voice cloning (consent-gated)","license":"bilibili Model Use License (commercial use allowed below 100M MAU and RMB 1B yearly revenue; may not be used to improve other models)","hf_repo":"IndexTeam/IndexTTS-2.5"}],"licence":"community","links":{"metrics":"/metrics/studio","page":"/studio","json":"/use-cases/studio.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"field","num":"05","name":"Field reports","status":"live","industries":["field-trades"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Expected issues in the report (grounded report + safety sweep + claim check)","value":"40/41 (98%); before 38/41 (93%)","unit":null,"n":41,"split":"synthetic","note":null},{"name":"Severity within the accepted range, of issues found","value":"40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)","unit":null,"n":40,"split":"synthetic","note":null},{"name":"Stated facts kept","value":"51/51; before 32/51 (63%)","unit":null,"n":51,"split":"synthetic","note":"incl. passed checks and spec limits"},{"name":"Claims the judge flagged as contradicted or unsupported","value":"5/125","unit":null,"n":125,"split":"synthetic","note":"2 are the overall-condition rating; 3 are speech-recognition spellings quoted verbatim. By hand: no invented findings."}],"dataset":"Internal synthetic eval: 8 TTS inspection walk-throughs (2 recorded demo + 6 new) with a gold checklist and a Gemma 4 judge, run through the live stack on decosa-api b24ab2a. The TTS audio was macOS system voices; this eval was not re-run after the demo audio was re-voiced with Decosa house voices (Kokoro-82M) on 26 Sep 2026.","held_out":false,"caveats":["The prompts were tuned on these same 8 scripts, so there is no held-out set.","Synthetic TTS audio only; no real site recordings.","Small n (8 walk-throughs).","Model judge (Gemma 4), checked by hand only in part."],"date":"2026-09-23","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · 4-bit on a 32 GB card, plus a small card for speech","evidence":[{"metric":"Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM","value":"478","source":"coding-agent-bench README (our server)"},{"metric":"Same weights served by llama.cpp (Q4 GGUF path): avoid for this tier","value":"62–63 / 500","source":"coding-agent-bench README (our server)"},{"metric":"Field lane quality","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · one 96 GB card (the hosted demo)","evidence":[{"metric":"Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM","value":"478","source":"coding-agent-bench README (our server)"},{"metric":"ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; FP8 on vLLM)","value":"34.2","source":"scribe-bench wiki/models.md"},{"metric":"Quality smoke, math / code","value":"38/40, 12/12","source":"Decosa model benchmarks (Sep 2026)"},{"metric":"Field report, internal synthetic eval (n = 8 walk-throughs: 2 recorded demo + 6 new), with the grounded report, safety sweep and claim check: expected issues in the report","value":"40/41 (98%); before 38/41 (93%)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: severity within the accepted range, of issues found","value":"40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: stated facts kept (incl. passed checks and spec limits, now fields of their own)","value":"51/51; before 32/51 (63%)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: missed safety item (bathroom receptacles without GFCI)","value":"in the report in 3 of 3 re-runs. Run on the original report, the safety sweep adds it as critical, quoted at 1:03","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: claims the judge flagged as contradicted or unsupported","value":"5/125 (before 0/100, plus 1 found by hand: 'cracked chimney' where the inspector said sealant). By hand: no invented findings. 2 of the 5 are the report's overall-condition rating; 3 are speech-recognition spellings now quoted verbatim ('ceiling' for sealant, 'pole' for pull stations, 'stab block' for Stab-Lok), which need an ASR glossary, not a model change","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"}]},{"tier":"best","label":"Best · two 96 GB cards","evidence":[{"metric":"ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; run on a Mac via MLX, not the NVFP4 kit)","value":"35.8","source":"scribe-bench wiki/models.md"},{"metric":"Medical-term miss rate on PriMock57, MOSS-Transcribe-Diarize vs Nemotron-3.5 streaming (%)","value":"8.4 vs 12.7","source":"scribe-bench RESULTS.md / wiki/decoder-finding.md"},{"metric":"Owner's coding benchmark, core tasks (/500), DeepSeek V4-Flash","value":"453–459","source":"coding-agent-bench README (our server)"},{"metric":"Field lane quality","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · the largest open flash models","evidence":[{"metric":"Field lane quality","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":26000,"p95_ms":null,"runs":null,"receipts_per_run":28,"cost_per_run_usd":0.026},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["The report is written after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load.","Measurements are checked against the spec the inspector states; there is no built-in code database."],"receipt_coverage":"full"},"cost_per_run_usd":0.026,"rehearsal_bundle":{"url":"/samples/field.zip","checks":9,"bytes":680859},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Speech recognition (streaming)","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Lanes and report (language model)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/field","page":"/tools/operations/field","json":"/use-cases/field.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"translate","num":"06","name":"Live translation","status":"live","industries":["general"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"FLORES-200 devtest en→es: chrF++ / BLEU / COMET-22","value":"55.3 / 29.9 / 87.2","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest es→en: chrF++ / BLEU / COMET-22","value":"59.7 / 31.8 / 87.7","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→fr: chrF++ / BLEU / COMET-22","value":"68.8 / 49.1 / 88.9","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→de: chrF++ / BLEU / COMET-22","value":"62.8 / 38.2 / 88.6","unit":null,"n":1012,"split":"heldout","note":null},{"name":"FLORES-200 devtest en→zh: chrF++ / BLEU / COMET-22","value":"29.7 / 45.7 / 89.5","unit":null,"n":1012,"split":"heldout","note":"zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese"}],"dataset":"FLORES-200 devtest (public, 1,012 sentences per direction) through the live NVFP4 endpoint with decosa-api's translate prompt.","held_out":true,"caveats":["Clean text rather than speech-recogniser output, so live-speech quality will be lower.","Standard tier only; lite, best and wanted tiers not measured yet.","Automatic metrics only; no human rating."],"date":"2026-09-23","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"Translation quality","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · one 96 GB card (hosted demo)","evidence":[{"metric":"FLORES-200 devtest, en→es / es→en (n = 1,012 sentences each): chrF++ / BLEU / COMET-22","value":"55.3 / 29.9 / 87.2 and 59.7 / 31.8 / 87.7","source":"; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)"},{"metric":"FLORES-200 devtest, en→fr / en→de / en→zh (n = 1,012 each): chrF++ / BLEU / COMET-22","value":"68.8 / 49.1 / 88.9; 62.8 / 38.2 / 88.6; 29.7 / 45.7 / 89.5 (zh BLEU with the zh tokenizer; chrF++ understates unsegmented Chinese)","source":"; eval results file translation-flores200-qwen38-27b.json (live NVFP4 endpoint, decosa-api's translate prompt, clean text rather than speech-recogniser output, 2026-09-23)"},{"metric":"Replay completeness, es→en and en→es","value":"28/28 sentences translated, 36/36 receipts signed","source":"decosa-api ops/record-all.json, our server 2026-09-23 (functional check, not a quality score)"},{"metric":"Engine quality smoke (math / code)","value":"38/40 and 12/12","source":"Decosa model benchmarks (Sep 2026) (pinned NVFP4+MTP3+FP8 KV stack)"}]},{"tier":"best","label":"Best · two 96 GB cards","evidence":[{"metric":"Translation quality","value":"not measured yet","source":null},{"metric":"Clinical note ROUGE-L on ACI-Bench (not translation; for relative ranking only)","value":"35.8 vs 34.2 for Qwen3.8-27B","source":"scribe-bench wiki/models.md (V4-Flash measured on a Mac Studio build, not this NVFP4 build)"}]},{"tier":"wanted","label":"Wanted · the largest open flash models","evidence":[{"metric":"Translation quality","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":2300,"p95_ms":null,"runs":null,"receipts_per_run":18,"cost_per_run_usd":0.0025},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 4 replay passed as written (14 translation lanes, glossary, summary, 18 attested receipts; a translation 157 ms median after its sentence), and the step 5 captions overlay translated live fake-microphone audio in headless Chromium once its createScriptProcessor buffer size was fixed. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["Only English and Spanish targets are tested; other targets are accepted but untested.","Before the merge, the hosted API refused lang=auto (the console's default), so Start recording on the translate page failed; samples worked."],"receipt_coverage":"full"},"cost_per_run_usd":0.0025,"rehearsal_bundle":{"url":"/samples/translate.zip","checks":8,"bytes":649621},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Speech recognition (streaming)","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Translation lane model (translation, glossary, summary)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/translate","page":"/tools/operations/translate","json":"/use-cases/translate.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"record","num":"07","name":"Tamper-evident record","status":"live","industries":["public-sector","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Claim check on the synthetic sessions: sentences supported","value":"14/17 council, 9/9 interview","unit":null,"n":26,"split":"synthetic","note":"Re-run 26 Sep 2026 after the sessions were re-voiced with Decosa house voices (Kokoro-82M). 2 of the 3 council flags come from one summary sentence split at \"Mr.\"; the first build (macOS voices): 19/19 and 9/9."},{"name":"Speaker attribution on the synthetic council meeting (5 TTS voices)","value":"5 speakers found; one short turn given to the wrong member","unit":null,"n":5,"split":"synthetic","note":"Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run; the first build (macOS voices): 4 speakers found, two male voices merged."}],"dataset":"Two synthetic TTS sessions (a council meeting and an interview) replayed through the live stack on a decosa-api test instance, one session at a time, gateway route.","held_out":false,"caveats":["WER on meeting or interview audio not measured yet; the ASR numbers published elsewhere are on clinical consultations (PriMock57).","Summary quality on meetings not measured yet.","Verifier recall on meeting minutes not measured yet.","Two synthetic sessions only."],"date":"2026-09-23","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card, captions only","evidence":[{"metric":"WER on meeting audio","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, two passes","evidence":[{"metric":"MOSS-TD WER / DER on PriMock57 (clinical proxy)","value":"10.3 / 11.4","source":"scribe-bench RESULTS.md"},{"metric":"claim check on the synthetic sessions: sentences supported","value":"19/19 and 9/9","source":"measured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route"},{"metric":"WER on meeting audio","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash writes and checks","evidence":[{"metric":"summary quality on meetings","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":16600,"p95_ms":null,"runs":null,"receipts_per_run":54,"cost_per_run_usd":0.0085},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 6 smoke passed as written: the record verified (132 entries), and after one edited word verification failed at the first edited transcript line. With the diarize service the record also carried speaker labels. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["Proves the record was not changed after signing and which key signed it; it does not prove the speech was recognised correctly. Not a certified court record.","Speaker labels need the diarize service; without it transcript lines have no speaker names."],"receipt_coverage":"full"},"cost_per_run_usd":0.0085,"rehearsal_bundle":{"url":"/samples/record.zip","checks":9,"bytes":645273},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Live captions (streaming, no speakers)","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"After the session: speaker-attributed transcript, one line per turn","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Actions lane, speaker roles, cited summary, claim verifier","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/record","page":"/tools/developer/record","json":"/use-cases/record.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"auditor","num":"09","name":"Endpoint auditor","status":"live","industries":["software","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b","value":"fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences","unit":null,"n":null,"split":"synthetic","note":"report aud_7be2d97ae80e20abdf34"},{"name":"Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)","value":"drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963","unit":null,"n":null,"split":"synthetic","note":"report aud_8a7c979db05198f582de"},{"name":"Reference noise, Qwen3.8-27B stack (6 runs)","value":"worst greedy repeat 5/10; mean |Δ logprob| ≤ 0.040; top-5 overlap ≥ 0.872","unit":null,"n":6,"split":"synthetic","note":"fixture qwen3.8-27b.json"},{"name":"Hosted gateway route vs reference","value":"pass; greedy 7/10 (band ≥ 4/10), canaries 8/10 = reference","unit":null,"n":null,"split":"synthetic","note":"report aud_687900df0710578a42a8"},{"name":"Raw engine route vs reference","value":"pass; greedy 10/10, logprobs identical over 244 tokens","unit":null,"n":null,"split":"synthetic","note":"report aud_b5a3508ddb34527f490e"},{"name":"Claim of BF16 weights (\"Qwen/Qwen3.8-27B\") when only the NVFP4 reference exists","value":"inconclusive, 4 of 4 claim runs (before 28 Sep it could be signed pass)","unit":null,"n":4,"split":"synthetic","note":"docs/evals/auditor-claims.md; our own direct engine"},{"name":"Genuine endpoint, correct claim, 28 Sep re-run","value":"hosted pass 3/3; direct engine pass 4/6 (2 drift in one run)","unit":null,"n":9,"split":"synthetic","note":"right after the production model restart; the noise band is too tight for a busy card"}],"dataset":"Audit reports and reference fixtures run on our server: two deliberate swaps (a smaller model and a re-quantised model under a false name), six reference runs to measure noise, and the hosted and raw routes against the reference.","held_out":false,"caveats":["Only two planted swaps, both set up by the builder; subtler substitutions were not tested.","A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.","On a busy card, a genuine self-hosted endpoint was flagged drift in about one run in three (1 of 3 on 24 Sep, 2 of 6 on 28 Sep), even though the re-check agreed on 28 Sep. Treat a single drift as a prompt to re-run, not a finding.","DeepSeek-V4-Flash reference fixture (best tier) not measured yet.","The claimed precision decides the reference: an official name such as Qwen/Qwen3.8-27B means BF16, and with no BF16 reference the verdict is inconclusive, not pass (28 Sep 2026)."],"date":"2026-09-24","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · audits only, no GPU","evidence":[{"metric":"Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b","value":"fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences","source":"report aud_7be2d97ae80e20abdf34, our server 2026-09-24"},{"metric":"Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)","value":"drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963","source":"report aud_8a7c979db05198f582de, our server 2026-09-24"}]},{"tier":"standard","label":"Standard · one 96 GB card (hosted demo)","evidence":[{"metric":"Reference noise, Qwen3.8-27B stack (16 runs, 12 of them under load)","value":"worst greedy repeat 4/10; mean |Δ logprob| ≤ 0.055; top-5 overlap ≥ 0.826","source":"fixture qwen3.8-27b.json, our server 2026-09-30 (re-recorded after the server was restarted with image and video input on 28 Sep)"},{"metric":"Hosted gateway route vs reference","value":"pass; greedy at or above the band (≥ 3/10), canaries equal to the reference. The nightly check runs this route","source":"the nightly check (its latest result is under 'How we tested it')"},{"metric":"Raw engine route vs reference","value":"pass in 10 of 10 audits on 30 Sep 2026; greedy 4 to 7 of 10 (band ≥ 3/10), top-5 overlap 0.835 to 0.893 (band ≥ 0.806)","source":"docs/evals/auditor-claims.md, 30 Sep 2026"},{"metric":"Engine drift on the same weights (coding benchmark, /500)","value":"478 NVFP4 + MTP; 387 FP8 eager; 193 FP8 + MTP; 62-63 llama.cpp CUDA","source":"coding-agent-bench README (not re-run by the auditor)"}]},{"tier":"best","label":"Best · two 96 GB cards","evidence":[{"metric":"DeepSeek-V4-Flash reference fixture","value":"not measured yet","source":"not measured yet"}]},{"tier":"wanted","label":"Wanted · references for the most-used open models","evidence":[{"metric":"substitution detection on these models","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":39822,"p95_ms":null,"runs":null,"receipts_per_run":22,"cost_per_run_usd":0.0018},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, audit of a local Qwen3.8-27B vLLM","notes":"Verified on 2026-09-25: signing key created, both reference fixtures listed, a full audit of a local Qwen3.8-27B vLLM (equivalent to the documented target) returned pass in 29 s with 23 probes, the report verified with the documented Python snippet and failed after an edit. The prompt's ./keys bind mount is not writable by the image user; the key was kept in the data volume instead (fix in progress)."},"known_limits":["Hosted figures are for the short audit (context probe off, 22 probes). The console's default run adds a ~12k-token context probe.","The hosted auditor only reaches public https:// endpoints; audit private or internal endpoints with the self-hosted auditor.","On a busy card, one self-hosted run in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it.","When the gateway is slow, the console's availability check marks the hosted target as not running and plays its recorded audit instead (seen 2026-09-25); POST /audit/runs still ran live.","A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.","Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"receipt_coverage":"full"},"cost_per_run_usd":0.0018,"rehearsal_bundle":{"url":"/samples/auditor.zip","checks":10,"bytes":16155},"models":[{"name":"decosa-api auditor (decosa_api/verticals/auditor)","role":"Probe runner, scorer and signer (no model; runs on CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Reference model (golden outputs, logprobs, noise band)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Qwen3.5-4B-Base (BF16)","role":"Small reference model (swap and quantisation demos)","license":"Apache-2.0","hf_repo":"Qwen/Qwen3.5-4B-Base"}],"licence":"permissive","links":{"metrics":"/metrics/auditor","page":"/tools/developer/auditor","json":"/use-cases/auditor.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"ugc","num":"15","name":"Disclosed UGC ads","status":"preview","industries":["sales-marketing","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Watermark survival on the two example ads (metadata stripped / CRF 28 re-encode / 50% resize)","value":"receipt id recovered in all 6 checks (6/7 to 7/7 frames vote)","unit":null,"n":6,"split":"synthetic","note":null},{"name":"Brand safety, automated check (RapidOCR rule) on 12 H3 clips","value":"flagged 1 of 2 clips a person rejected, 0 of 10 approved clips","unit":null,"n":12,"split":"dev","note":"calibration set"}],"dataset":"Two example ads for watermark robustness, and 12 H3 clips reviewed by a person to calibrate the brand-lettering check (our server, 2026-09-24).","held_out":false,"caveats":["Ad quality not measured with a metric; the lite tier was judged not usable for presenter ads by owner review.","The brand check was calibrated on the same 12 clips it is scored on, and it missed 1 of 2 rejected clips.","Lip-sync quality not measured yet."],"date":"2026-09-24","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · fully open (Apache-2.0), on your own GPU","evidence":[{"metric":"Ad quality","value":"judged not usable next to H3 or LTX-2.3 for presenter ads (owner review, 25 Sep 2026); not measured with a metric","source":"owner review"},{"metric":"Render time","value":"788 s per 5 s shot at FP8 weights, 30 steps (one RTX PRO 6000)","source":"measured on our server 2026-09-23"}]},{"tier":"standard","label":"Standard · the hosted default, MiniMax H3 via fal at cost","evidence":[{"metric":"Watermark survival on the two example ads (metadata stripped / CRF 28 re-encode / 50% resize)","value":"receipt id recovered in all 6 checks (6/7 to 7/7 frames vote)","source":"measured on our server 2026-09-24, <data>/ugc/watermark-robustness-h3.json"},{"metric":"Brand safety, automated check (RapidOCR rule) on 12 H3 clips","value":"flagged 1 of 2 clips a person rejected, 0 of 10 approved clips","source":"calibration 2026-09-24, decosa_api/verticals/ugc/brand_check.py"},{"metric":"Ad quality","value":"not measured yet","source":null}]},{"tier":"best","label":"Best for self-host · MiniMax H3 (licence pending) or LTX-2.3 on your own GPU","evidence":[{"metric":"Lip-sync and ad quality","value":"not measured yet in this product","source":null}]},{"tier":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","evidence":[{"metric":"render time per clip against one card with offload","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"partial","p50_ms":6695,"p95_ms":null,"runs":null,"receipts_per_run":6,"cost_per_run_usd":0.0027},"selfhost":{"date":"2026-09-25","result":"partial","method":"Fresh clone of decosa-api, api image plus the prompt's extras, compose from the assemble prompt with llm and comfy pointed at an already-running Qwen3.8-27B vLLM and ComfyUI on the same box.","notes":"Verified on 2026-09-25: images build, the service starts, /ugc/policy, the refused brief (no model call), a plain brief (seven lanes, six receipts signed by the box's key, status attested, 7.6 s) and pricing all work against local model servers equivalent to the documented ones; model-server startup was not re-verified. The Kokoro voice-over helper ran on CPU in the container (2 s line in 6 s). A Wan2.1 ad render (about 13 minutes per shot on the GPU) was not run. Fixed on the way: the image left out services/ (first render failed with 'tts exited 2'); self-hosted boxes defaulted to the fal render path; Kokoro's spaCy model download. Worked around locally (fixed in the shared self-host pass): data folders owned by root, and a curl health check the image cannot run."},"known_limits":["The same brief can plan with render_allowed true on one run and false on the next: the writer is a model. The server revises a script once; after that, plan again or add evidence.","Hosted renders need the fal account to have credit; while it is empty, plans still work and renders are refused up front with a clear message.","Hosted plan latency depends on the shared model server: about 7 s when quiet, 16-40 s when busy.","Frames are checked for brand lettering by OCR, not by a person; the published example ads were also reviewed by a person."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0027,"rehearsal_bundle":{"url":"/samples/ugc.zip","checks":11,"bytes":1950},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Lane model (brief check, hooks, script, claims review, disclosure text, storyboard)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MiniMax H3 Max (via fal)","role":"Hosted default video: presenter still, lip-synced presenter shots and reference-consistent B-roll (with audio)","license":"H3's own licence excludes the US; used here only through fal's hosted endpoints, listed as commercial use under fal's MiniMax partnership","hf_repo":null},{"name":"Kokoro-82M","role":"Voice-over (preset synthetic voices only)","license":"Apache-2.0","hf_repo":"hexgrad/Kokoro-82M"},{"name":"c2pa-rs via c2pa-python","role":"Content credential and render receipt (provenance kit)","license":"MIT OR Apache-2.0","hf_repo":null},{"name":"TrustMark variant Q","role":"Invisible watermark on every frame (payload: the receipt id)","license":"MIT","hf_repo":null}],"licence":"community","links":{"metrics":"/metrics/ugc","page":"/studio/brand","json":"/use-cases/ugc.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"grounding","num":"17","name":"Grounding check","status":"live","industries":["general","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Unsupported-sentence precision / recall, default gate","value":"0.593 / 0.556","unit":null,"n":2069,"split":"heldout","note":"F1 0.574, agreement 0.929, Cohen's kappa 0.535, flag rate 8.1%."},{"name":"Recall, strict gate (partial counts too)","value":"0.966","unit":null,"n":2069,"split":"heldout","note":"Precision 0.276, flag rate 30.1%."},{"name":"Response-level F1, default gate","value":"0.691","unit":null,"n":300,"split":"heldout","note":"P 0.706, R 0.675."},{"name":"Baseline: NLI cross-encoder on CPU (lite tier), F1","value":"0.227","unit":null,"n":2069,"split":"heldout","note":"P 0.133, R 0.792, flag rate 51.4%."},{"name":"Default gate precision / recall on dev (tuned on)","value":"0.488 / 0.494","unit":null,"n":927,"split":"dev","note":null}],"dataset":"RAGTruth (MIT): human span labels on responses from six LLMs to QA, news summaries and data-to-text; dev 120 responses from train, test 300 responses (2,069 sentences, 178 unsupported) from the test split, run once after freezing.","held_out":true,"caveats":["One benchmark, English only, generated by 2023-era models; the false-flag rate must be measured per domain before the gate blocks without a human look.","A date line was added to the prompt after the dev run and before the test run.","The eval ran on the direct route to the same model server, so its calls carry no gateway receipts.","RAGTruth annotators are lenient on added detail, so some strict-gate flags count against the checker without being wrong; the labels were not re-annotated.","'Supported by the sources' is not 'true': a sentence copied from a wrong source passes."],"date":"2026-09-24","doc_url":"https://decosa.ai/metrics/evals/grounding"},"quality_evidence":[{"tier":"lite","label":"Lite · CPU only, no GPU","evidence":[{"metric":"RAGTruth test, sentence level: precision / recall on unsupported","value":"0.133 / 0.792 (F1 0.227)","source":"docs/evals/grounding.md, 300 held-out responses, threshold from dev"},{"metric":"RAGTruth test: agreement with human labels","value":"53.6% (κ 0.09)","source":"docs/evals/grounding.md"},{"metric":"Lexical-overlap baseline, same test","value":"0.177 / 0.455 (F1 0.255), agreement 77.1%","source":"docs/evals/grounding.md"}]},{"tier":"standard","label":"Standard · one GPU for the judge (hosted demo)","evidence":[{"metric":"RAGTruth test, default gate (block = unsupported or contradicted): precision / recall on unsupported","value":"0.593 / 0.556 (F1 0.574)","source":"docs/evals/grounding.md, 2,069 sentences, 178 unsupported, held out"},{"metric":"RAGTruth test, default gate: agreement with human labels","value":"92.9% (κ 0.535)","source":"docs/evals/grounding.md"},{"metric":"RAGTruth test, strict (partial also flagged): precision / recall","value":"0.276 / 0.966 (F1 0.429), agreement 77.9%","source":"docs/evals/grounding.md"},{"metric":"RAGTruth test, response level, default gate: precision / recall","value":"0.706 / 0.675","source":"docs/evals/grounding.md"}]},{"tier":"wanted","label":"Wanted · a panel of the largest open judges","evidence":[{"metric":"sentence-level F1 on the grounding set, same protocol as standard","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":17202,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0021},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers","notes":"Verified on 2026-09-25, option A: the api starts, the thermostat sample blocks with the battery and 240 V sentences contradicted and the reset sentence supported at S1.3, the stream and the signed report behave as documented (p50 1.9 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Option B (CPU NLI judge) also ran: it blocked the sample but marked the battery-life sentence supported."},"known_limits":["The judge reads only the sources given: \"supported\" is not \"true\".","Each sentence is one model call that re-sends the sources, so tokens grow with sentences x source length (about 5.8k for a five-sentence answer against a one-page manual).","The CPU NLI judge (self-host option B) is coarser: on the thermostat sample it missed one of the two contradictions.","Sources must be pasted text; a bare URL is refused (nothing is fetched).","Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"receipt_coverage":"full"},"cost_per_run_usd":0.0021,"rehearsal_bundle":{"url":"/samples/grounding.zip","checks":10,"bytes":2105},"models":[{"name":"decosa-api grounding module (decosa_api/verticals/grounding)","role":"Checker: segmentation, evidence selection, gate, signed report (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Judge: one call per sentence, verdict and cited spans","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/grounding","page":"/tools/developer/grounding","json":"/use-cases/grounding.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"deposition","num":"18","name":"Deposition and hearing digest","status":"live","industries":["legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Contradiction finder: planted conflicts found (held-out set)","value":"9 / 9 in both runs","unit":null,"n":9,"split":"test","note":null},{"name":"Contradiction finder: false-positive flags, run 1 / run 2","value":"2 / 1 (precision 0.82 / 0.90)","unit":null,"n":5,"split":"test","note":"All false flags were planted traps; 15 flags over 2 runs, very small"},{"name":"Digest sentences fully supported by their cited lines (hand-checked)","value":"33 / 40","unit":null,"n":40,"split":"dev","note":"7 partly supported (mostly a dropped hedge), 0 not supported; one annotator, not a lawyer"},{"name":"Cite checker: wrong cite flagged","value":"59 (100%)","unit":null,"n":59,"split":"synthetic","note":null},{"name":"Cite checker: changed fact flagged","value":"49 (91%)","unit":null,"n":54,"split":"synthetic","note":"Unchanged sentences flagged on a second check: 3 of 60 (5%)"}],"dataset":"Two public-domain congressional hearing excerpts (govinfo), a fictional two-witness demo pair used while writing the prompts, and a held-out fictional contradiction set (3 matters, 9 planted conflicts, 5 traps, 1 ambiguous pair).","held_out":true,"caveats":["The held-out set was written by the same agent that wrote the prompts, before the eval ran.","Nothing was checked by a lawyer, and no real deposition was used.","Coverage (does the digest include everything important) is not measured; a 300-page deposition is not measured.","The audio path is tested only with a fake diarizer; its accuracy on deposition audio is unknown.","A number guard added after seeing the misses (51/54) is measured on the set it was designed from, so it is not held out."],"date":"2026-09-24","doc_url":"https://decosa.ai/metrics/evals/deposition"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"cite-check accuracy on transcripts","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"hand check: sentences fully supported by their cited lines","value":"33/40","source":"decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; checked by the building agent, not a lawyer"},{"metric":"cite checker: wrong cites / changed facts flagged","value":"59/59 / 49 of 54","source":"decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route"},{"metric":"contradictions, held-out fictional set: recall / precision","value":"9/9 / 0.82-0.90 (two runs)","source":"decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route"},{"metric":"cite check on real deposition transcripts (by a lawyer)","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"digest and check quality","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · GLM-5.3-Flash on your own hardware","evidence":[{"metric":"digest quality","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":40081,"p95_ms":44103,"runs":5,"receipts_per_run":48,"cost_per_run_usd":0.013804},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 5 smoke passed as written in 23 s: every lane, 54 attested receipts, the 2:15 pm / 11:30 am conflict marked INCONSISTENT, the record verified and the Word export opened. A PDF transcript parsed as numbered (4:1-6:14), and with the diarize service /deposition/rough returned an uncertified rough transcript. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","The claim check shows each sentence is supported by the lines it cites; it does not show the digest covers everything important.","Hosted is for public-record and fictional transcripts only."],"receipt_coverage":"full"},"cost_per_run_usd":0.013804,"rehearsal_bundle":{"url":"/samples/deposition.zip","checks":9,"bytes":3930},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Digest writer, cite checker and contradiction judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/deposition","page":"/legal/deposition","json":"/use-cases/deposition.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"filing-preflight","num":"21","name":"Filing pre-flight","status":"live","industries":["legal"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Fake citations caught as problem (strict)","value":"89%","unit":null,"n":19,"split":"test","note":"17 of 19; 95% lenient (problem or review). Wilson 95% 0.69-0.97."},{"name":"Citations for a holding the case does not contain, caught","value":"92%","unit":null,"n":13,"split":"test","note":"12 of 13, strict and lenient."},{"name":"One-word misquotations caught as problem (strict)","value":"64%","unit":null,"n":22,"split":"test","note":"95% lenient; the 7 'review' ones are singular/plural changes."},{"name":"False problems on 9 unaltered real briefs","value":"19 of 511 (3.7%)","unit":null,"n":511,"split":"test","note":"Plus 67 items (13%) marked review; Wilson 95% 0.02-0.06."},{"name":"Privacy pattern findings on 12 real briefs","value":"0","unit":null,"n":12,"split":"test","note":"570,911 characters of public-domain briefs."}],"dataset":"12 real Solicitor General briefs from 2025-2026 (public domain; 3 dev, 9 test, first ~9,000 characters of the argument) with planted errors: Mata v. Avianca fake cites, moved pages, hand-written wrong propositions and one-word misquotes.","held_out":true,"caveats":["Small: 12 briefs and 54 planted errors on test, so the rates have wide intervals.","One author wrote the wrong propositions, and they lean toward clear reversals; subtle mischaracterisations are harder.","The real briefs come from one careful filer; briefs citing more unpublished or Westlaw-only decisions are covered less well.","Propositions are a triage list: 47% of real-brief propositions came back 'review'.","No human cite-checker was timed against it, and there is no good-law (citator) signal."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/filing-preflight"},"quality_evidence":[{"tier":"lite","label":"Lite · CPU only, no GPU","evidence":[{"metric":"Fake citations caught (existence and name checks need no model)","value":"17/19 problem, same as standard","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs; the existence check is identical without the judge"},{"metric":"One-word misquotations (string match, no model)","value":"14/22 problem, 21/22 problem or review, same as standard","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"Citations for a holding the case does not contain","value":"not checked in this tier","source":"no judge"}]},{"tier":"standard","label":"Standard · one GPU for the judge (hosted demo)","evidence":[{"metric":"Fake citations caught (the six Mata v. Avianca fakes plus real cites with the first page moved): problem / problem or review","value":"17/19 (89%) / 18/19 (95%)","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"Citations for a holding the case does not contain, caught as problem","value":"12/13 (92%)","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"One-word misquotations: problem / problem or review","value":"14/22 (64%) / 21/22 (95%); the 7 reviews are singular/plural changes","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"False problems on the same briefs unaltered (511 checked items)","value":"19 (3.7%): quotations tied to the wrong source 11, authorities the free indexes lack 6, name parse 1, proposition 1","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"Propositions read by the judge on real briefs: supported / partly / not supported by passages read / contradicted","value":"29 / 24 / 25 / 4 of 82 (all but 1 shown as review, not problem)","source":"docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs"},{"metric":"Privacy patterns on the full text of 12 real briefs (570,911 characters)","value":"0 false findings","source":"docs/evals/filing-preflight.md"}]},{"tier":"wanted","label":"Wanted · a GLM-5.3-Flash judge on your own hardware","evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":12063,"p95_ms":19008,"runs":5,"receipts_per_run":15,"cost_per_run_usd":0.015458},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Offline (the default in the prompt) the planted record-cite and privacy problems are found and the cases come back unverified, as documented; with lookups on, the same 9 problems as hosted. The signed record verifies and a tampered entry is named."},"known_limits":["No citator: it does not say whether a case is still good law. Westlaw- or Lexis-only decisions, many unpublished orders, agency decisions and state codes come back unverified, never OK.","Hosted speed depends on load on the shared service: the time shown is the median of our latest production runs of the fictional sample."],"receipt_coverage":"full"},"cost_per_run_usd":0.015458,"rehearsal_bundle":{"url":"/samples/filing-preflight.zip","checks":12,"bytes":5739},"models":[{"name":"decosa-api filing pre-flight (decosa_api/verticals/preflight)","role":"Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Judge: one call per proposition and per record cite, plus one call for minors' names","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/filing-preflight","page":"/legal/filing-preflight","json":"/use-cases/filing-preflight.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"promo-claims-check","num":"22","name":"Promotional-claims pre-check","status":"live","industries":["healthcare","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted problems caught (all categories)","value":"20 / 20","unit":null,"n":20,"split":"test","note":"Repeat run of the same test split: 20 / 20. Dev: 11 / 11."},{"name":"Piece-level problems caught (fair balance, DSHEA)","value":"2 / 2","unit":null,"n":2,"split":"test","note":null},{"name":"Unplanted sentences with an issue (false positives)","value":"1 / 63 (1.6%)","unit":null,"n":63,"split":"test","note":"Repeat run: 1 / 63. Dev: 2 / 31."},{"name":"Unplanted sentences with any finding (issue or check)","value":"3 / 63","unit":null,"n":63,"split":"test","note":"Repeat run: 4 / 63."},{"name":"Compliant pieces with no issue","value":"3 / 4","unit":null,"n":4,"split":"test","note":null}],"dataset":"12 synthetic promotional pieces for a fictional distributor (4 dev, 8 test), checked against public FDA labels (openFDA), NIH Office of Dietary Supplements fact sheets and synthetic spec sheets. Each product has one piece with planted problems and one written to comply.","held_out":true,"caveats":["The planted problems are blatant and were written by the same person who built the checker; subtle problems are not measured.","Synthetic copy only; the false-positive rate is on copy written to comply, and real copy should bring more check-level noise.","Text only: visual prominence of risk information is not measured, and devices were not evaluated.","Greedy decoding on a shared server is not bit-for-bit repeatable; one recording run flagged a compliant piece.","One change was made after the first dev run; the test split was never used to change anything."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/promo-claims-check"},"quality_evidence":[{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Planted problems caught, held-out test (unsupported, off-label, comparative, disease)","value":"20 / 20 (9/9, 3/3, 3/3, 5/5), same in a repeat run","source":"docs/evals/promo-claims-check.md: 8 synthetic pieces on 3 public labels and 3 NIH fact sheets"},{"metric":"Piece-level problems caught, test (fair balance, DSHEA disclaimer)","value":"2 / 2","source":"docs/evals/promo-claims-check.md"},{"metric":"False positives, test: unplanted sentences with an issue / with any finding","value":"1 / 63 (1.6%) / 3-4 of 63","source":"docs/evals/promo-claims-check.md, two runs"},{"metric":"Dev split (the only split used for a change)","value":"11 / 11 caught; 2 / 31 unplanted with an issue","source":"docs/evals/promo-claims-check.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":3917,"p95_ms":null,"runs":null,"receipts_per_run":14,"cost_per_run_usd":0.0225},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Planted metformin piece: status issues with every expected finding (contradicted HbA1c, off-label weight and heart claims, comparative, fair balance, boxed warning); the compliant piece came back clean; the packet verifies and a changed finding fails."},"known_limits":["Speed depends on load: a 7-14 sentence piece took 4-10 s on a quiet GPU and 30-35 s while the shared GPU was busy (25 Sep 2026).","Text only: type size, placement and contrast are not assessed. Tables flattened to text can be misread.","The eval's planted problems are blatant ones; subtle violations are not measured yet. A person reviews every finding."],"receipt_coverage":"full"},"cost_per_run_usd":0.0225,"rehearsal_bundle":{"url":"/samples/promo-claims-check.zip","checks":10,"bytes":5529},"models":[{"name":"decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding)","role":"Pre-check: sentences, label sections, rules, findings, the packet (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Grounding judge, claim reviewer, head-to-head and fair-balance checks","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/promo-claims-check","page":"/tools/life-sciences/promo-claims-check","json":"/use-cases/promo-claims-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"migration-check","num":"23","name":"Open-model migration check","status":"live","industries":["software"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Not-worse agreement with human experts, both orders (run 1)","value":"79.3% (74.4-83.5)","unit":null,"n":300,"split":"test","note":"Runs 2 and 3: 77.3% and 78.3%. GPT-4 (MT-Bench's own judge) on the same items: 78.0%. On par, not better."},{"name":"Cohen's kappa vs experts (run 1)","value":"0.588","unit":null,"n":300,"split":"test","note":"Runs 2 and 3: 0.546, 0.568; GPT-4: 0.562"},{"name":"Precision / recall on \"worse\" (run 1)","value":"0.755 / 0.839","unit":null,"n":300,"split":"test","note":"Errs on the strict side more than the lenient one."},{"name":"Verdict identical in all three runs, per item","value":"84.3%","unit":null,"n":300,"split":"test","note":"Temperature 0 on a batched server is not bit-exact; variation lands on close calls."},{"name":"Report-level verdict equal to the experts' labels","value":"14 of 15","unit":"pairings","n":15,"split":"test","note":"GPT-4: 13 of 15. 14 of 15 pairings kept the same verdict across three runs."},{"name":"Dev agreement, both orders","value":"85.0%","unit":null,"n":120,"split":"dev","note":"GPT-4 87.5% on the same dev items"}],"dataset":"lmsys/mt_bench_human_judgments (CC-BY-4.0): expert pairwise votes on first-turn MT-Bench answers from six 2023-era models, with GPT-4's own verdicts as the baseline judge. Split by question: 120 dev items, 300 test items; the prompt was written once, run once on dev, not changed, and test was run three times.","held_out":true,"caveats":["Measures only the free-text judge; the structured scorers are deterministic code covered by unit tests.","MT-Bench answers are 2023-era and general-purpose; agreement on a team's own task should be checked against a few of their own labels.","Human tie votes are noisy (about a quarter of items); not-worse folds them into \"not worse\".","First-turn answers only; no multi-turn conversations.","Self-preference when the judge model judges its own answers is not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/migration-check"},"quality_evidence":[{"tier":"lite","label":"Lite · structured prompts on a 24-32 GB card (self-host)","evidence":[{"metric":"Structured scoring (exact, label, JSON schema, fields)","value":"deterministic code, covered by unit tests; no model to measure","source":"tests/test_migration.py"},{"metric":"Gemma-4-31B as a candidate","value":"not measured yet","source":"not measured yet"}]},{"tier":"standard","label":"Standard · Qwen3.8-27B as candidate and judge (hosted demo)","evidence":[{"metric":"MT-Bench test, agreement with human experts on 'is the candidate worse?' (both orders)","value":"79.3% (95% CI 74.4-83.5), κ 0.588; runs 2 and 3: 77.3%, 78.3%","source":"docs/evals/migration-check.md, 300 held-out pairs"},{"metric":"Same items, GPT-4 as judge (published MT-Bench verdicts)","value":"78.0% (73.0-82.3), κ 0.562","source":"docs/evals/migration-check.md"},{"metric":"Agreement without ties (judge vs experts)","value":"89.5% (GPT-4: 88.0%)","source":"docs/evals/migration-check.md"},{"metric":"Report verdict unchanged across 3 runs / matches the experts' verdict","value":"14 of 15 pairings / 14 of 15 (GPT-4: 13 of 15)","source":"docs/evals/migration-check.md"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host)","evidence":[{"metric":"As a migration candidate","value":"not measured yet","source":"not measured yet"}]},{"tier":"wanted","label":"Wanted · the largest open candidates","evidence":[{"metric":"agreement with expert labels, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":3603,"p95_ms":null,"runs":null,"receipts_per_run":30,"cost_per_run_usd":0.0029},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. 20 synthetic tickets: schema valid 20 of 20, a verdict with reasons; the record verifies and recomputes, and a flipped score fails at that entry with the numbers that no longer follow named. Key minting with the admin secret works."},"known_limits":["The hosted numbers are for the 20-ticket JSON sample. The 20-question MT-Bench free-text sample (70 calls with the judge) took 19 s on a quiet GPU and 3-7 minutes while the shared GPU was busy (25 Sep 2026).","Agreement with your current model is not correctness; send human labels where you have them.","Latency in the report is measured on a shared GPU through the gateway; your own deployment will differ."],"receipt_coverage":"full"},"cost_per_run_usd":0.0029,"rehearsal_bundle":{"url":"/samples/migration-check.zip","checks":10,"bytes":3510},"models":[{"name":"decosa-api migration module (decosa_api/verticals/migration)","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Candidate and free-text judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/migration-check","page":"/tools/developer/migration-check","json":"/use-cases/migration-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"typed-judgment","num":"24","name":"Typed-judgment API","status":"live","industries":["general","software"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"BoolQ accuracy / ECE, logprobs (direct route)","value":"90.7% / 0.023","unit":null,"n":1000,"split":"heldout","note":"AUROC 0.857. Samples k=4 (hosted default): ECE 0.023, AUROC 0.651."},{"name":"MMLU accuracy / ECE, logprobs (direct route)","value":"83.4% / 0.026","unit":null,"n":1000,"split":"heldout","note":"AUROC 0.863. Samples k=4: ECE 0.053, AUROC 0.768."},{"name":"Hosted route, samples k=4: BoolQ / MMLU accuracy","value":"90.5% / 84.0%","unit":null,"n":200,"split":"heldout","note":"200 questions each; ECE 0.032 / 0.028; $0.61 / $0.58 per 1,000."},{"name":"SummEval rubric, mean Spearman vs expert mean (logprobs)","value":"0.525","unit":null,"n":25,"split":"heldout","note":"25 held-out articles, 1,600 judgments."},{"name":"MT-Bench pairwise agreement with expert votes, with ties (logprobs)","value":"64.8%","unit":null,"n":600,"split":"heldout","note":"81.3% without ties; GPT-4 pair judge on the same rows 65.2%. Pairwise probabilities are poorly calibrated (ECE 0.17)."},{"name":"Same answer when the temperature-0 call is repeated, BoolQ / MMLU (direct route)","value":"99.0% / 95.7%","unit":null,"n":1000,"split":"heldout","note":null}],"dataset":"Public benchmarks with fixed dev/test splits (seed 24): BoolQ (300 dev, 1,000 test), MMLU (300 dev, 1,000 test), SummEval (10 dev, 25 test articles), MT-Bench human judgments (150 dev, 600 test votes). Calibration fit on dev only; each test split run once.","held_out":true,"caveats":["Logprobs, the only method that ranks right against wrong answers well, need the direct route (self-host) today; the hosted API uses samples.","Stated confidence is almost always 95-100 and carries little information.","A busy inference server is not bit-reproducible: repeated temperature-0 calls change some answers, mostly near ties.","Only these public benchmarks: calibration varies by domain, so check on your own labelled data. Multi-label use was not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/typed-judgment"},"quality_evidence":[{"tier":"lite","label":"Lite · one call per question (stated confidence)","evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)","value":"90.7% / 0.076 / 0.088 / 0.703","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)","value":"83.4% / 0.124 / 0.152 / 0.602","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"}]},{"tier":"standard","label":"Standard · hosted, answer plus 4 seeded samples","evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)","value":"90.7% / 0.023 / 0.079 / 0.651","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)","value":"83.4% / 0.053 / 0.114 / 0.768","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"BoolQ, 200 questions through the hosted gateway route: accuracy / ECE / Brier","value":"90.5% / 0.032 / 0.077","source":"docs/evals/typed-judgment.md (service run)"},{"metric":"MMLU, 200 questions through the hosted gateway route: accuracy / ECE / Brier","value":"84.0% / 0.028 / 0.121","source":"docs/evals/typed-judgment.md (service run)"},{"metric":"MT-Bench pairwise, samples k=4: agreement with ties / without ties","value":"64.5% / 81.0%","source":"docs/evals/typed-judgment.md"}]},{"tier":"best","label":"Best · self-host with logprobs","evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)","value":"90.7% / 0.023 / 0.069 / 0.857","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)","value":"83.4% / 0.026 / 0.102 / 0.863","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"SummEval rubric (25 held-out articles): mean per-article Spearman with experts, expected score","value":"0.525 (logprobs) · 0.491 (samples, hosted)","source":"docs/evals/typed-judgment.md; G-Eval with GPT-4 reported 0.514 on the full set (different protocol)"},{"metric":"MT-Bench pairwise (600 held-out expert votes): agreement with ties / without ties","value":"64.8% / 81.3%; GPT-4 judge on the same rows 65.2% / 83.6%","source":"docs/evals/typed-judgment.md; GPT-4 verdicts from the dataset's gpt4_pair split"}]},{"tier":"wanted","label":"Wanted · two large judges that must agree","evidence":[{"metric":"BoolQ and MMLU accuracy and calibration, same splits as standard","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":4126,"p95_ms":null,"runs":null,"receipts_per_run":35,"cost_per_run_usd":0.0044},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. \"method\": \"auto\" used logprobs as documented (7 calls for the ticket); two runs gave the same verdicts hash, probabilities moved by up to 0.02; the record verifies as signed by this box and a changed answer fails."},"known_limits":["Speed depends on load: the support-ticket sample (4 questions, 4 samples each, 35 calls) took about 4 s on a quiet GPU and 25-80 s while the shared GPU was busy (25 Sep 2026).","Log-probabilities, the best-calibrated method, are self-host only until the gateway passes them through; hosted requests use samples or stated confidence.","Calibration was measured on public benchmarks and varies by domain: check it on your own labelled data before a threshold decides anything."],"receipt_coverage":"full"},"cost_per_run_usd":0.0044,"rehearsal_bundle":{"url":"/samples/typed-judgment.zip","checks":10,"bytes":2204},"models":[{"name":"decosa-api judgment module (decosa_api/verticals/judgment)","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Judge: one temperature-0 call per question, plus seeded samples","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/typed-judgment","page":"/tools/developer/typed-judgment","json":"/use-cases/typed-judgment.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"security-questionnaire","num":"25","name":"Security questionnaire answerer","status":"live","industries":["compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Fill precision","value":"0.970 (64 of 66)","unit":null,"n":66,"split":"test","note":"Dev: 0.963 (26 of 27). BM25 baseline on test: 0.606. With candidates from the evidence retrieval block (27 Sep, same test split): 0.984 (63 of 64)."},{"name":"Coverage (answerable questions filled correctly)","value":"0.984 (63 of 64)","unit":null,"n":64,"split":"test","note":"BM25 baseline: 0.672."},{"name":"Abstention (unanswerable questions sent to a person)","value":"0.969 (31 of 32)","unit":null,"n":32,"split":"test","note":"Dev: 0.917 (11 of 12). BM25 baseline: 0.625. With the evidence retrieval block (27 Sep, same test split): 1.000 (32 of 32)."},{"name":"Stale approved answers caught","value":"4 of 4","unit":null,"n":4,"split":"test","note":null},{"name":"Filled answers with an invented claim","value":"0","unit":null,"n":66,"split":"test","note":"Lexical check: 0 misses."}],"dataset":"A fictional vendor's library (50 approved answers, 9 documents; 2 answers deliberately out of date) and hand-written questions labelled fill, flag or stale before any model run: 40 dev, 100 test held out and run once after the prompts were frozen.","held_out":true,"caveats":["Synthetic, single-author, one small library, English only: the same author wrote the library, the questions and the labels.","Question phrasing follows the library's vocabulary more closely than real buyer sheets; expect lower coverage and more review on real questionnaires.","Omissions are not checked: a trimmed answer can drop a qualifying sentence and still pass.","Only two stale entries, a very small sample.","\"Supported\" is a model judgement; the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/security-questionnaire"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card, no policy check","evidence":[{"metric":"Held-out test: fill precision without the source check","value":"0.928 (64 of 69)","source":"derived from docs/evals/security-questionnaire/test.json: the 3 answers the policy check sent to review would have been filled"},{"metric":"Held-out test: coverage / abstention","value":"0.984 / 0.969 (unchanged: the policy check only affects stale answers)","source":"docs/evals/security-questionnaire.md"}]},{"tier":"standard","label":"Standard · one GPU, selection plus both checks (hosted demo)","evidence":[{"metric":"Held-out test (100 questions): fill precision","value":"0.970 (64 of 66)","source":"docs/evals/security-questionnaire.md, synthetic library, labels written before any run"},{"metric":"Held-out test: coverage of answerable questions","value":"0.984 (63 of 64)","source":"docs/evals/security-questionnaire.md"},{"metric":"Held-out test: correct abstention when nothing approved fits","value":"0.969 (31 of 32)","source":"docs/evals/security-questionnaire.md"},{"metric":"Held-out test: stale approved answers caught","value":"4 of 4","source":"docs/evals/security-questionnaire.md"},{"metric":"Invented claims in filled answers (dev + test, 93 fills)","value":"0; lexical check 0 misses","source":"docs/evals/security-questionnaire.md"},{"metric":"Baseline, BM25 top-1 with a threshold from dev: precision / coverage / abstention","value":"0.606 / 0.672 / 0.625","source":"docs/evals/security-questionnaire.md"},{"metric":"With the evidence retrieval block (27 Sep, same held-out test, run once): precision / coverage / abstention / stale","value":"0.984 (63 of 64) / 0.984 (63 of 64) / 1.000 (32 of 32) / 4 of 4","source":"docs/evals/retrieval.md; runs in docs/evals/security-questionnaire/retrieval/"},{"metric":"Prompt tokens for the 100 held-out questions, BM25 candidates vs the retrieval block","value":"218,203 vs 151,150 (-31%)","source":"docs/evals/retrieval.md"},{"metric":"Retrieval alone (reranker top-1, threshold from dev, no model call): precision / coverage / abstention","value":"0.864 / 0.797 / 0.844 (BM25: 0.606 / 0.672 / 0.625)","source":"docs/evals/retrieval.md"}]},{"tier":"wanted","label":"Wanted · the largest open models, long context","evidence":[{"metric":"answer precision against the BM25 baseline, same protocol as standard","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":62952,"p95_ms":null,"runs":null,"receipts_per_run":62,"cost_per_run_usd":0.0313},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. The 30-question sample: 21 approved, 2 review (TVM-01, LOG-01), 7 for a person, in 19 s with 61 calls; the XLSX export has a row per question with blank answers on the none rows; the record verifies and a changed status fails."},"known_limits":["Speed depends on load: the 30-question sample took about 20 s on a quiet self-hosted GPU, about 60 s hosted, and about 2-3 minutes while the shared GPU was busy (25 Sep 2026).","PDFs and Word files are not read: send their text as documents.","The checks are model judgements and do not catch an answer that leaves out a qualifier; a person reads the sheet before it is sent. The eval library is synthetic."],"receipt_coverage":"full"},"cost_per_run_usd":0.0313,"rehearsal_bundle":{"url":"/samples/security-questionnaire.zip","checks":10,"bytes":6795},"models":[{"name":"decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks","role":"Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","role":"Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per question","license":"Apache-2.0 (both)","hf_repo":"Qwen/Qwen3-Reranker-4B"},{"name":"Qwen3.8-27B (NVFP4)","role":"Selector (one call per question) and grounding judge (one call per changed or checked sentence)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/security-questionnaire","page":"/tools/finance/security-questionnaire","json":"/use-cases/security-questionnaire.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"flight-recorder","num":"26","name":"Agent flight recorder","status":"live","industries":["software","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Genuine records that verify","value":"19/19","unit":null,"n":19,"split":"synthetic","note":"3 hosted-demo, 8 held-out SDK runs, 7 imported Jev runs, 1 synthetic OpenAI-shaped loop."},{"name":"Tampered copies caught, issuer key pinned","value":"339/339","unit":null,"n":339,"split":"synthetic","note":"22 kinds of alteration."},{"name":"Tampered copies caught without key pinning","value":"312/339","unit":null,"n":339,"split":"synthetic","note":"The 27 misses are chains rebuilt and re-signed with another key, which only pinning catches (by design)."},{"name":"Held-out agent runs reaching the expected outcome","value":"8/8","unit":null,"n":8,"split":"heldout","note":"4 tasks x 2 repeats; not a benchmark."},{"name":"MiniWoB++ success, production agent with guards","value":"25.9%","unit":null,"n":625,"split":"heldout","note":"95% CI 22.6-29.5%; 35.2% without the guards; no tuning for the bench."},{"name":"Mind2Web element accuracy / step success","value":"44.7% / 40.7%","unit":null,"n":300,"split":"heldout","note":"Raw model answers; 35% step success under the production value guard."},{"name":"Recording overhead per step with a screenshot (median)","value":"18.3","unit":"ms","n":40,"split":"synthetic","note":"8.4 ms hash-only."}],"dataset":"19 sealed flight records (hosted demo, held-out SDK runs on saucedemo.com and a fictional shop, imported Jev macOS runs) with 339 tamper trials; plus the demo agent on MiniWoB++ (125 tasks x 5 seeds) and a 300-step Mind2Web sample.","held_out":true,"caveats":["The record proves what was reported, not what happened: a client-reported step is only as honest as the agent.","Without key pinning, a record rebuilt and re-signed with another key verifies (0 of 27 caught).","Agent success of 8/8 on four held-out tasks is small; the decision prompt was adjusted on the three hosted demo tasks.","On public benchmarks the agent is demo-grade: 25.9% on MiniWoB++, failing on canvas, drag, custom widgets and unquoted values.","MiniWoB++ is tiny and synthetic and Mind2Web is offline; no live multi-page benchmark (WebArena) was run.","Decision latency was measured under heavy shared load."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/flight-recorder"},"quality_evidence":[{"tier":"lite","label":"Lite · record only, any CPU","evidence":[{"metric":"Genuine records that verify","value":"19/19 (3 hosted demo runs, 8 SDK agent runs, 7 imported Jev harness runs, 1 computer-use loop)","source":"decosa-api docs/evals/flight-recorder.md, 2026-09-25"},{"metric":"Tampered copies caught (22 kinds of alteration)","value":"339/339 with the issuer key pinned; 312/339 without (the 27 were rebuilt and re-signed with another key, which only pinning can catch)","source":"decosa-api docs/evals/flight-recorder-results.json"},{"metric":"Overhead per step","value":"18 ms and 15 KB of record with a thumbnail; 8 ms and 3.3 KB hash-only","source":"decosa-api docs/evals/flight-recorder.md (40-step benchmark, local HTTP)"}]},{"tier":"standard","label":"Standard · receipted decisions, one 96 GB card (hosted demo)","evidence":[{"metric":"Held-out agent tasks reaching the expected outcome (2 on saucedemo.com, 2 on the demo shop, 2 runs each; 2 expected a guard stop)","value":"8/8","source":"decosa-api docs/evals/flight-recorder.md; the prompt was adjusted on the three demo tasks, not these"},{"metric":"Decisions with a gateway-signed receipt","value":"83/83","source":"decosa-api docs/evals/flight-recorder.md"},{"metric":"Guard stops on an order or payment button","value":"3/3 runs whose task asked to place or finish an order stopped before the click","source":"decosa-api docs/evals/flight-recorder.md"},{"metric":"MiniWoB++ success, the agent on its own (125 tasks x 5 seeds, BrowserGym task classes)","value":"25.9% (95% CI 22.6-29.5%) with the guards; 35.2% (31.6-39.0%) without them; 36.7% on the form-and-button tasks, 4% on canvas, drag and slider tasks","source":"decosa-api docs/evals/computer-use-bench.md, 2026-09-25; production prompt, no tuning"},{"metric":"Mind2Web, next-step accuracy on real websites (300 test steps, top-50 candidates)","value":"element 44.7% (39.1-50.3%), step success 40.7% (35.3-46.3%) before the guards; 35% after them","source":"decosa-api docs/evals/computer-use-bench.md, 2026-09-25"}]}],"benchmark":{"title":"How well does the agent do on its own?","intro":"The recorder's numbers above measure the record. This measures the demo agent: Qwen3.8-27B choosing actions from the numbered element table, on two public benchmarks, with the production prompt and no tuning. 25 Sep 2026, 95% intervals in brackets.","rows":[{"label":"MiniWoB++, 125 small web tasks x 5 seeds, production agent with guards","value":"25.9% (22.6-29.5%)","detail":"36.7% on tasks built from links, buttons and inputs; 4% on canvas, shape and colour tasks; 4% on drag, slider and keyboard tasks"},{"label":"Same, without the guards","value":"35.2% (31.6-39.0%)","detail":"The 9-point gap is almost all the value guard: it refused to type dates, sums and other values the task did not quote"},{"label":"Same weights reading the screenshot instead of the table (not served today)","value":"55.8% (51.9-59.7%)","detail":"Pixels only, coordinate clicks, no guards; 0.75 s per decision on a private server"},{"label":"Mind2Web, 300 recorded steps on real websites: right element / right element and operation","value":"44.7% / 40.7%","detail":"Brackets 39.1-50.3% and 35.3-46.3%. Under the production value guard, 35% of steps would go through"},{"label":"Decision time, production agent","value":"0.59 s median","detail":"p90 7.1 s when the shared model server was busy"}],"points":[{"heading":"What the element table means","text":"The model never sees the page. It gets a numbered list of the links, buttons, inputs and selects it can act on, and picks one. That makes each choice checkable and recordable, and it is why it does well on ordinary forms. It is also a crutch: anything not in the list does not exist for the agent. In 20% of MiniWoB episodes the list was empty at the first step, because the clickable things were plain spans and divs."},{"heading":"Where it fails","text":"Canvas and drawn shapes, colours, drag and sliders, custom widgets such as date pickers, and content inside iframes, which the extractor does not enter. Long, exploratory sequences fail too: an eight-step flight booking scored 0 of 5 with every setup, and paging through tabs to find a link often loops until the step limit. In most of these cases it says it is blocked rather than guessing."},{"heading":"Why the guards matter","text":"Without them, the model said it was done in 82 of 625 runs (13%) when the task had not finished, and it typed values nobody gave it. With them, a run is only marked successful when the page agrees, and anything typed comes from the task or the caller's allowed values. That costs capability, and the table shows how much. Pass the values a task needs as allowed_values rather than turning the guard off."},{"heading":"Verdict","text":"Demo-grade on its own. It is usable for narrow, form-shaped jobs on sites with real links and inputs, when the values are supplied and a completion check is written for the task, as in the hosted demo and the 8/8 held-out runs. It is not a general web agent. The vision result shows where the gain is: a hybrid that uses the table when it has the target and the screenshot when it does not, with the same guards."}],"source":"decosa-api docs/evals/computer-use-bench.md and computer-use-bench.json (every episode and model answer), 2026-09-25. MiniWoB++ (MIT) through BrowserGym (Apache-2.0); Mind2Web (CC BY 4.0) with the Multimodal-Mind2Web test subset (OpenRAIL). WebArena was not run: it needs six self-hosted sites."},"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":5293,"p95_ms":null,"runs":null,"receipts_per_run":8,"cost_per_run_usd":0.0025},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api (ba02fab), api image built from docker/api/Dockerfile (705 MB, no browser), compose from the assemble prompt with the llm service dropped and DECOSA_LLM_URL pointed at an already-running Qwen3.8-27B vLLM on the same box.","notes":"Verified on 2026-09-25: the image builds, the service starts, and the smoke test passes end to end against a local model server equivalent to the documented one; model-server startup itself was not re-verified. Key minting, the prompt's smoke script (verify, then fail at step 1 after a change), /decide (click on element 1, receipt status attested, 238 ms), 20-step timing (18.5 ms per step with a 480 KB screenshot, 3.9 ms hash-only) and the site's record viewer pointed at the box all worked. Worked around locally (fixed in the shared self-host pass): the compose health check calls curl, which the image does not have."},"known_limits":["Hosted demo sessions are limited per network each hour (the current number is in GET /healthz); the console also takes an API key.","A key has a per-minute request rate and a cap on open runs (GET /flight/info lists the limits). Close an abandoned run with DELETE /flight/runs/<id>, or see your runs with GET /flight/runs; the Python SDK waits out a 429.","Hosted decision latency depends on the shared model server: under a second per decision when quiet, several seconds when busy.","A restart of the hosted service stops a demo run in progress (shown as 'the demo run stopped').","The record proves what was reported and that it was not changed after signing; it does not prove a website did what it showed."],"receipt_coverage":"full"},"cost_per_run_usd":0.0025,"rehearsal_bundle":{"url":"/samples/flight-recorder.zip","checks":9,"bytes":39879},"models":[{"name":"decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK","role":"Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)","license":"AGPL-3.0-or-later (the SDK, the flight record format and its verifier are Apache-2.0)","hf_repo":null},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Decision model: picks the next action from a numbered element table (never coordinates, never free text)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Playwright 1.58 with Chromium headless shell","role":"Headless browser for the hosted demo and the Playwright adapter","license":"Apache-2.0 (Playwright); BSD-3-Clause (Chromium)","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/flight-recorder","page":"/tools/developer/flight-recorder","json":"/use-cases/flight-recorder.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"test-runs","num":"27","name":"Verified end-to-end test runs","status":"live","industries":["software","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Verdict agrees with the scripted Playwright test, first run of each case","value":"52/52","unit":null,"n":52,"split":"test","note":null},{"name":"Verdict agrees, all runs","value":"120/120","unit":null,"n":120,"split":"test","note":null},{"name":"Seeded bugs and faulty users caught","value":"24/24","unit":null,"n":24,"split":"test","note":"Each at the same step as the scripted test."},{"name":"False fails on working builds (including the redesign)","value":"0/96","unit":null,"n":96,"split":"test","note":null},{"name":"Flaky cases (verdict changed between repeats)","value":"0/52","unit":null,"n":52,"split":"test","note":"Every case ran 2-4 times; bounds the flip rate only loosely (roughly under 3% at 95% confidence)."},{"name":"Tampered certificates caught","value":"1,320/1,320","unit":null,"n":1320,"split":"test","note":null}],"dataset":"52 cases: 5 specs x 8 builds of a fixture shop app written for this eval (a good build, a redesign and 6 seeded bugs), plus 3 saucedemo.com specs x 4 public test users with known faults. Ground truth from hand-written Playwright scripts; 120 runs; 11 tamper alterations on every certificate.","held_out":true,"caveats":["Small, and written by us: the fixture app, its bugs and the specs were written by the same author as the runner (the saucedemo faults are Sauce Labs').","Nothing was tuned on these cases (the action prompt is the flight recorder's, unchanged), but 52/52 shows it works on simple shop flows, not on a large product.","The model is language-only: canvas-heavy and cross-origin-iframe apps are not covered, and layout or visual regressions are out of scope.","Latency was measured while the shared GPU was saturated by other evals."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/test-runs"},"quality_evidence":[{"tier":"lite","label":"Lite · navigation-only specs, any CPU","evidence":[{"metric":"Navigation steps in the eval","value":"86 goto steps, all judged by the same assertion code as the agent steps","source":"decosa-api docs/evals/test-runs.md"}]},{"tier":"standard","label":"Standard · agent steps, one 96 GB card (hosted demo)","evidence":[{"metric":"Verdict agrees with a hand-written Playwright test (52 cases: 5 specs x 8 fixture builds, 3 saucedemo specs x 4 users)","value":"52/52 first runs; 120/120 over all runs","source":"decosa-api docs/evals/test-runs.md, 2026-09-25"},{"metric":"Seeded bugs and faulty users caught","value":"24/24 runs, each at the same step as the scripted test; 0/96 false fails on working builds, including a UI redesign","source":"decosa-api docs/evals/test-runs-results.json"},{"metric":"Flaky cases over 2-4 repeats","value":"0/52","source":"decosa-api docs/evals/test-runs.md"},{"metric":"Tampered certificates caught (11 alterations)","value":"1,320/1,320 with the issuer key pinned","source":"decosa-api docs/evals/test-runs-results.json"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":86054,"p95_ms":null,"runs":null,"receipts_per_run":6,"cost_per_run_usd":0.0015},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, image built with WITH_BROWSER=1, compose up, sample against local model servers","notes":"Images build, the service starts, the sample passes on the good build (exit 0, 9 s) and fails on the seeded bug (exit 1, 19 s) against the already-running local Qwen3.8-27B vLLM (host network, no llm service started); a flipped assertion fails /testruns/verify; a private http target listed in DECOSA_TESTRUNS_TARGETS passes and an unlisted one is refused. Named volume for /data. Model-server startup itself not re-verified."},"known_limits":["Eval cases are simple shop flows written for it (plus Sauce Labs' public site); expect more stuck steps on complex apps.","The action model reads the element table, not pixels: canvas-heavy and cross-origin iframe apps are not supported.","Assertions read the DOM; layout and visual regressions are out of scope.","A run takes about 1.5 minutes when the shared GPU is busy (seconds when it is quiet); a scripted Playwright test is faster and free.","Hosted runs only reach the fixture app, saucedemo.com and domains you verified; hosted certificates are kept 24 hours."],"receipt_coverage":"full"},"cost_per_run_usd":0.0015,"rehearsal_bundle":{"url":"/samples/test-runs.zip","checks":8,"bytes":1611},"models":[{"name":"decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client","role":"Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)","license":"AGPL-3.0-or-later (the decosa_testrun CI client is Apache-2.0)","hf_repo":null},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Playwright 1.58 with Chromium headless shell","role":"Headless browser that carries out the steps and reads the assertions","license":"Apache-2.0 (Playwright); BSD-3-Clause (Chromium)","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/test-runs","page":"/tools/developer/test-runs","json":"/use-cases/test-runs.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"model-risk-pack","num":"28","name":"Model-risk evidence pack","status":"live","industries":["finance","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"False alarms on unchanged runs (ALERT / WATCH)","value":"0 of 8 / 0 of 8","unit":null,"n":8,"split":"test","note":"95% interval for the alert rate on 8 runs: 0-32%. Dev: 0 of 4."},{"name":"Injected changes detected (ALERT)","value":"8 of 12","unit":"runs","n":12,"split":"test","note":"8 of 8 for prompt, bias and model changes; each ALERT named the right lane."},{"name":"Sampling drift detected (temperature 0.3 / 0.7)","value":"0 of 4","unit":"runs","n":4,"split":"test","note":"1 WATCH, not reproduced on re-check."},{"name":"Adverse-action drafts lane separating injections","value":"5/5 in every configuration","unit":null,"n":null,"split":"test","note":"The drafts grader did not separate any injection: weak evidence as built."},{"name":"Cost per unchanged pack (list price)","value":"$0.013","unit":null,"n":null,"split":"test","note":"82 calls; a pack that re-checks: $0.02-0.03"}],"dataset":"Synthetic suite fernhill-v1: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action drafts, plus 10 golden prompts. Dev (2 unchanged runs per system, 4 injections not reused) set the policy; test (8 unchanged runs, 6 different injections x 2 runs) was run once after the policy was fixed.","held_out":true,"caveats":["Synthetic suite and injections written by the same team that built the pack.","Small sample: 8 unchanged runs gives a 0-32% interval on the false-alarm rate.","Sampling or settings drift on a black box is not detected.","Fairness probes only see the attributes and values in the suite; a bias on an unprobed value would pass.","The drafts lane grader checks presence, not specificity."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/model-risk-pack"},"quality_evidence":[{"tier":"lite","label":"Lite · grader on one 32 GB card (self-host)","evidence":[{"metric":"Detection and false alarms on this hardware","value":"not measured yet","source":"not measured yet"}]},{"tier":"standard","label":"Standard · hosted grader and demo systems (Qwen3.8-27B)","evidence":[{"metric":"False alarms on unchanged systems (held-out)","value":"0 of 8 ALERT, 0 of 8 WATCH (95% CI for the alert rate 0-32%)","source":"docs/evals/model-risk-pack.md"},{"metric":"Injected prompt, bias and model changes caught (held-out)","value":"8 of 8 ALERT: vendor prompt update 2/2, ZIP-code bias 2/2, age bias 2/2, swap to Qwen3-1.7B 2/2; each named the right lane","source":"docs/evals/model-risk-pack.md"},{"metric":"Sampling drift caught (temperature 0.3 and 0.7)","value":"0 of 4 ALERT (1 WATCH): the triage decisions did not change, so the pack did not alarm","source":"docs/evals/model-risk-pack.md"},{"metric":"Adverse-action drafts lane","value":"5/5 in every configuration: it did not separate these injections (weak evidence as built)","source":"docs/evals/model-risk-pack.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":62000,"p95_ms":null,"runs":null,"receipts_per_run":82,"cost_per_run_usd":0.013},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the pre-release branch, api image built from docker/api/Dockerfile, compose from the assemble prompt (llm service dropped, api on host network pointed at the running Qwen3.8-27B vLLM, named volume).","notes":"Verified on 2026-09-25: image builds, service starts healthy, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Lender pack PASS (15/15, 5/5, 82 attested receipts, 12 s); /mrm/verify ok, and a one-word edit in the record fails at that entry; key minting and two mrm_monitor.py runs (baseline, then trend) worked; logs held no case text. Torn down afterwards."},"known_limits":["Detects only what the suite probes: other products, attributes or ZIP codes are not covered.","Sampling or settings changes that do not change decisions are not detected (0 of 4 in the eval).","The drafts grader checks that listed reasons appear, not how specific they are.","Hosted demo sessions allow about three packs (20,000 generated tokens); use an API key for more.","Scheduled monitoring is a client script (cron or systemd), not a hosted scheduler."],"receipt_coverage":"full"},"cost_per_run_usd":0.013,"rehearsal_bundle":{"url":"/samples/model-risk-pack.zip","checks":10,"bytes":2925},"models":[{"name":"decosa-api model-risk module (decosa_api/verticals/mrm)","role":"Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Fixed grader (typed judgments) and the hosted system under test","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Qwen3-1.7B (BF16, CPU)","role":"Injected problem in the eval: the 'vendor' silently moved to a small model","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-1.7B"}],"licence":"permissive","links":{"metrics":"/metrics/model-risk-pack","page":"/tools/finance/model-risk-pack","json":"/use-cases/model-risk-pack.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"clinical-ai-monitor","num":"29","name":"Clinical AI assurance monitor","status":"live","industries":["healthcare","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Invented fact caught as an error","value":"49 (98%)","unit":null,"n":50,"split":"test","note":null},{"name":"Changed detail caught as an error","value":"47 (94%)","unit":null,"n":50,"split":"test","note":"With the detail checker (M17). The judge alone: 36 (72%). False error flags on faithful sentences unchanged: 4 of 1,135"},{"name":"Key item left out caught","value":"45 (90%)","unit":null,"n":50,"split":"test","note":"48/50 after a guard was removed post-test; that number is not held out"},{"name":"Wrong speaker caught / typed correctly","value":"47 (94%) / 31 (62%)","unit":null,"n":50,"split":"test","note":null},{"name":"Invented exam finding caught as an error","value":"49 (98%)","unit":null,"n":50,"split":"test","note":null},{"name":"False error flags on faithful notes (sentences)","value":"4 (0.35%)","unit":null,"n":1135,"split":"test","note":"Key items called left out: 3 of 368 (0.8%); notes with at least one error finding: 6 of 50 (12%)"},{"name":"Tighten: words removed (median per note)","value":"15%","unit":null,"n":50,"split":"test","note":"Range 0-31%; 15,835 to 13,499 words over 50 clean notes; 0 words added (checked in code on every output)"},{"name":"Tighten: key items that stopped being fully recorded","value":"5 of 362 (1.4%)","unit":null,"n":362,"split":"test","note":"All 5 came back 'partly recorded', none missing; e.g. 'over the counter' dropped from a medicine"},{"name":"Changed detail caught as an error, 200 more blind plants","value":"188 (94.0%)","unit":null,"n":200,"split":"test","note":"95% CI 89.8-96.5%; judge alone 147 (73.5%); written by a blind sub-agent on the 50 held-out visits, never trained on"},{"name":"Detail checker alone on CPU: changed details caught / faithful sentences flagged","value":"33 of 50 (66%) / 3 of 1,135 (0.26%)","unit":null,"n":50,"split":"test","note":"No language model, windows by word overlap, p(changed) >= 0.9; median 7.1 s per note on 8 CPU threads"}],"dataset":"All 57 PriMock57 mock primary-care consultations (CC BY 4.0) with reference transcripts; one synthetic cited scribe note per visit plus five copies each with one planted error; 7 visits dev, 50 held out.","held_out":true,"caveats":["The notes are synthetic and the plants are clean single errors; real scribe errors are subtler, so these detection rates are an upper bound.","Reference transcripts: no speech-recognition error enters this eval.","One judge model family; a second judge is not run. Plants were checked by a validator, not reviewed by a clinician.","Primary care, English, remote visits, 50 test visits; no specialty or inpatient data and no clinician agreement study.","The detail checker's thresholds were fixed on the dev split before the test run. On ACI-Bench's human notes (real speech-recognition transcripts) it turned 9 of 1,469 sentences into errors: 2 real conflicts, 3 details never said aloud, 4 not errors. On ACI-Bench plants the judge alone caught 71/78 and with the checker 72/78.","A blind frontier model (Claude Opus via Claude Code) proofreading 40 of the same notes caught 20/20 changed details with 1 flag on 20 faithful notes; this tool caught 19/20 with 2 flags, on open weights you can run yourself."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/clinical-ai-monitor"},"quality_evidence":[{"tier":"lite","label":"Lite · text transcripts, one 32 GB card","evidence":[{"metric":"Same judge and prompts as standard, so the same detection and false-flag rates","value":"see standard","source":"docs/evals/clinical-ai-monitor.md"}]},{"tier":"standard","label":"Standard · judge plus diarizer (hosted demo)","evidence":[{"metric":"Planted errors caught at error severity (50 held-out PriMock57 visits, one error per note): invented fact / invented exam finding / wrong speaker / key item left out / changed detail","value":"49/50 / 49/50 / 47/50 / 45/50 / changed detail 47/50 with the detail checker (36/50 judge alone); every plant flagged at least as review (50/50 each)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Changed details, 200 more blind plants on the same 50 held-out visits (8 detail types)","value":"188/200 (94%) with the detail checker, 147/200 (73.5%) judge alone; false error flags on faithful notes unchanged (4/1,135)","source":"docs/evals/clinical-ai-monitor.md (M17)"},{"metric":"Wrong-speaker errors typed as wrong speaker","value":"31/50 (62%); the rest flagged as contradicts or not in the visit","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Error flags on faithful notes (1,135 sentences, 368 key items)","value":"4 sentences (0.35%; 1 real error in the note, 2 judge mistakes, 1 debatable), 3 key items (0.8%; 1 real)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Clinicians' own PriMock57 notes: lines with an error finding","value":"12.1%; in a sample of 20, 16 were statements the transcript does not support (names, routine negatives never asked)","source":"docs/evals/clinical-ai-monitor.md"},{"metric":"Speech recognition, if you send audio (MOSS-Transcribe-Diarize, PriMock57)","value":"10.3% WER","source":"scribe-bench RESULTS.md (the clinical scribe's pass 2)"}]},{"tier":"wanted","label":"Wanted · a second judge from another family","evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":28000,"p95_ms":41400,"runs":null,"receipts_per_run":30,"cost_per_run_usd":0.021},"selfhost":{"date":"2026-09-25","result":"pass","method":"assemble-prompt.md on our server: fresh clone into a clean directory, api image built from docker/api/Dockerfile, compose with a named volume, pointed at the already-running Qwen3.8-27B vLLM (127.0.0.1:8114) and MOSS diarizer (127.0.0.1:8092) instead of starting new ones; then torn down","notes":"Images build, the service starts, and the smoke steps pass: the wrong-speaker sample flagged (17 s), the report verifies with transcript and note hashes, a changed count fails verification, the summary signs, and 170 s of audio came back as 22 speaker turns. Model-server startup itself was not re-verified. /healthz says ok: false on this stack because it also checks the live-scribe recogniser, which the monitor does not use."},"known_limits":["Measured on synthetic notes with one clear planted error each; real scribe errors are subtler.","Changed details: 94% caught as errors with the detail checker (47/50 and 188/200 blind plants), 72-74% by the judge alone; the checker only upgrades the judge's own 'detail not in the visit' flags.","On real speech-recognition transcripts the detail checker adds some false error flags (4 in 1,469 ACI-Bench sentences, e.g. 'type i' vs 'type 1'); each comes with its transcript line.","Wrong-speaker errors are caught but typed correctly only 62% of the time.","Tighten shortened the 50 test notes by a median 15%; 5 of 362 recorded key items came back 'partly recorded' after tightening (none missing).","Measures against the transcript it is given; speech-recognition errors pass into the measure.","Clinicians' own notes get flagged for things never said aloud; use it on AI drafts.","Audio input is API-only (POST /monitor/transcribe); the page takes text."],"receipt_coverage":"full"},"cost_per_run_usd":0.021,"rehearsal_bundle":{"url":"/samples/clinical-ai-monitor.zip","checks":10,"bytes":8074},"models":[{"name":"decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)","role":"Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)","role":"Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5","license":"Apache-2.0","hf_repo":"decosaai/decosa-note-detail-checker-modernbert-large"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Audio in (optional): one-pass speaker-attributed transcript of the whole visit","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/clinical-ai-monitor","page":"/clinics/clinical-ai-monitor","json":"/use-cases/clinical-ai-monitor.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"privilege-log","num":"30","name":"Privilege review and privilege log","status":"live","industries":["legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Privileged vs not, model call: accuracy / precision / recall","value":"83.9% / 82.3% / 79.7%","unit":null,"n":149,"split":"test","note":"Direct route with logprobs. Clear labels (92): 95.7%; hard labels (57): 64.9%. Dev: 95.9%."},{"name":"AUROC of p(privileged) / ECE","value":"0.908 / 0.058","unit":null,"n":149,"split":"test","note":null},{"name":"Privileged emails the tool would have produced (waiver risk)","value":"4 (6.2% of privileged)","unit":null,"n":64,"split":"test","note":"All four on labels marked hard. Wrongly withheld: 3."},{"name":"Sent to attorney review","value":"57 (38%)","unit":null,"n":149,"split":"test","note":"Auto-decided documents agreeing with the labels: 85 of 92 (92.4%)."},{"name":"Four-way call: exact / Cohen's kappa","value":"80.5% / 0.65","unit":null,"n":149,"split":"test","note":"Kappa 0.83 on clear labels"},{"name":"Planted leaky log descriptions caught","value":"19 of 21","unit":null,"n":37,"split":"synthetic","note":"0 false alarms on 16 clean descriptions; code rules alone 12 of 21."}],"dataset":"200 emails from Enron lawyers' mailboxes in the public corpus (198 labelled; 49 dev, 149 test), a 28-email synthetic set and 37 planted log descriptions. Dev was used to write the prompts and review rules; test was run once after the rules were frozen, with nothing changed because of it.","held_out":true,"caveats":["Labels are one AI reviewer's (Claude), not a lawyer's; 72 of 198 are marked hard judgment calls.","Small sample; publish your own measured precision and recall on a labelled sample of the matter before relying on it.","The held-out Enron set was run on the direct (self-host, logprobs) route only; the hosted gateway route was measured only on the 28 synthetic emails.","On hard calls the model alone is barely better than chance (kappa 0.28); the review queue, not the model, is the safeguard."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/privilege-log"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card, self-hosted","evidence":[{"metric":"Calls on held-out Enron email","value":"not measured separately: the same weights and prompts as the standard tier, so the calls should match; speed on a 5090 not measured","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Privileged vs not, model call, 149 held-out Enron emails (64 privileged by our labels): accuracy / precision / recall / AUROC of p(privileged)","value":"83.9% / 82.3% / 79.7% / 0.908","source":"docs/evals/privilege-log.md, direct route with logprobs; labels written by Claude (an AI agent), not a lawyer"},{"metric":"After the review rules: privileged emails the tool would have produced / non-privileged it would have withheld / sent to attorney review","value":"4 of 64 (6.2%) / 3 of 85 / 57 of 149 (38%); decided without review: 92, of which 92.4% agree with the labels","source":"docs/evals/privilege-log.md, same run"},{"metric":"The same, on the 92 held-out emails whose label we marked clear (not a judgment call)","value":"accuracy 95.7%, 0 privileged produced, 0 wrongly withheld, 26% to review","source":"docs/evals/privilege-log.md; all 4 misses and 3 over-withholds were on the 57 emails we marked as hard calls"},{"metric":"Four-way call (attorney-client / work product / both / not privileged), held out","value":"80.5% exact, Cohen's kappa 0.65","source":"docs/evals/privilege-log.md"},{"metric":"Planted leaky log descriptions caught (21 leaky, 16 clean, synthetic documents)","value":"19 of 21 caught, 0 of 16 false alarms (code rules alone 12 of 21; the judge catches paraphrases)","source":"docs/evals/privilege-log.md, leak-check stage, direct route"},{"metric":"Drafted log descriptions that leaked (87 on held-out Enron, 18 on the synthetic set)","value":"1 first draft flagged (a subject phrase copied from the email), rewritten; 0 in the final log by the code rules","source":"docs/evals/privilege-log.md"},{"metric":"Synthetic set (28 emails): labels met / expected flags and groups met / consistency across copies, drafts and threads / same decision on a re-run","value":"0 privileged produced, 0 wrongly withheld, 3 to review / 12 of 12 / 7 of 7 groups consistent / 28 of 28 (27 of 28 same privilege type)","source":"docs/evals/privilege-log.md"},{"metric":"Hosted gateway route on the synthetic set (sampled call): privileged produced / wrongly withheld / to review / expectations met","value":"0 / 0 / 3 / 12 of 12; same privileged-vs-not accuracy as the direct route (96.4%), AUROC 0.974","source":"docs/evals/privilege-log.md, hosted route stage"}]},{"tier":"wanted","label":"Wanted · a larger second judge on your own hardware","evidence":[{"metric":"accuracy and AUROC, same protocol as standard","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":121000,"p95_ms":null,"runs":null,"receipts_per_run":100,"cost_per_run_usd":0.031},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs, the harborline-dispute sample passes end to end (copy of HL-002 gets its call, HL-003 goes to review, HL-012 and HL-021 produced, 17 s, 58 attested calls), the signed record verifies and fails when one call is changed, and the CSV export works. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Labels are one AI reviewer's (Claude's), not a lawyer's; hard judgment calls are where it errs: 4 of 34 hard privileged held-out emails would have been produced.","About 38% of held-out real email goes to attorney review (26% where the label is clear).","The hosted gateway route slows sharply when the shared gateway is loaded (one 12-email run took 938 s).","Attachments must be sent as separate documents; no OCR, PDF or native file parsing in this version.","Partial privilege (redacting part of a document) is not proposed; such documents go to review."],"receipt_coverage":"full"},"cost_per_run_usd":0.031,"rehearsal_bundle":{"url":"/samples/privilege-log.zip","checks":11,"bytes":2805},"models":[{"name":"decosa-api privilege module (decosa_api/verticals/privilege)","role":"Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/privilege-log","page":"/legal/privilege-log","json":"/use-cases/privilege-log.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"editorial-ledger","num":"31","name":"Editorial-control ledger","status":"live","industries":["creative-media","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Edit metrics exact against scripted diffs","value":"80 / 80","unit":null,"n":80,"split":"synthetic","note":"First run 59 / 80; a harness bug and a line-wrap bug were fixed before the second run."},{"name":"Rubber-stamped vs heavily edited separated (AUC, words changed)","value":"1.00","unit":null,"n":43,"split":"synthetic","note":"32 rubber-stamped vs 11 heavy edits; the heavy edits are the builder's (3) and the model's (8), not newsroom editors'."},{"name":"Wording-only paraphrases with no claim change flagged (held out)","value":"24 / 24","unit":null,"n":24,"split":"heldout","note":"Generated after the last prompt change and never used to tune it."},{"name":"Claim-change flag on planted edits: recall / specificity","value":"24/24, 40/40","unit":null,"n":64,"split":"synthetic","note":"The style-only row was seen during prompt tuning, so it is optimistic."},{"name":"IteraTeR test, meaning-changed edits flagged (recall)","value":"31 of 35 (0.89)","unit":null,"n":35,"split":"test","note":"Precision against the intent label 0.49 (32 of 65 other edits flagged); the label is a proxy, not 'a claim changed'."},{"name":"Tampered ledgers caught","value":"120 / 120","unit":null,"n":120,"split":"synthetic","note":"15 kinds of change on 8 ledgers; all 8 genuine ledgers verify."}],"dataset":"8 fictional newsroom source packs with a Qwen3.8-27B draft each, scripted and planted edits on them, plus the IteraTeR human revisions (Apache-2.0; dev split for prompt work, test split for the numbers).","held_out":true,"caveats":["No real newsroom editors on real copy: the heavy edits are the builder's and the model's.","The style-only claim-diff cases were seen while tuning the prompt; only the 24 paraphrases are a fair wording-only number.","Claim-diff precision was not measured on a set labelled for claim changes; IteraTeR intent labels are a proxy.","A careful editor who changes nothing looks the same as a rubber stamp; the ledger does not judge review quality.","Someone holding the server's signing key who also forges a fresh gateway receipt is not caught by the record alone.","Pieces longer than about 24,000 characters and quiet-GPU latency are not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/editorial-ledger"},"quality_evidence":[{"tier":"lite","label":"Lite · ledger only, on CPU (self-host)","evidence":[{"metric":"Edit metrics against scripted diffs of real drafts","value":"80 / 80 exact","source":"docs/evals/editorial-ledger.md"},{"metric":"Rubber-stamped vs heavily edited, words changed","value":"AUC 1.00; rubber-stamped at most 0.68%, heavy at least 51.1%","source":"docs/evals/editorial-ledger.md (32 rubber-stamped, 11 heavy)"},{"metric":"Tampered ledgers caught","value":"120 / 120 (15 kinds of change, 8 ledgers); 8 / 8 genuine verified","source":"docs/evals/editorial-ledger.md"}]},{"tier":"standard","label":"Standard · Qwen3.8-27B drafts and checks claims (hosted demo)","evidence":[{"metric":"Claim check on planted edits (number changed, fact deleted or added, paragraphs moved, style-only, 24 held-out paraphrases)","value":"64 / 64; recall 24/24, no false alarm in 40","source":"docs/evals/editorial-ledger.md"},{"metric":"Claim check on IteraTeR human sentence edits (test)","value":"31 / 35 meaning-changed caught; fired on 32 / 65 others (precision 0.49 against the intent label, which undercounts real fact changes)","source":"docs/evals/editorial-ledger.md"},{"metric":"Drafts with a gateway-signed receipt covering the stored text","value":"8 / 8 in the eval; 2 / 2 model calls per run in the end-to-end runs","source":"docs/evals/editorial-ledger.md; scripts/smoke/editorial-ledger.py"},{"metric":"Edit metrics, separation and tamper detection","value":"as the lite tier (same code)","source":"docs/evals/editorial-ledger.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass (pre-release server)","p50_ms":58653,"p95_ms":null,"runs":null,"receipts_per_run":2,"cost_per_run_usd":0.002},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory on our server, image built from docker/api/Dockerfile, api started with compose (named volume, python healthcheck), then the assemble prompt's smoke steps 1-8 and the CMS webhook with a minted dk_ key; torn down afterwards.","notes":"Verified 25 Sep 2026: image builds, the service starts healthy, and the sample passes end to end against local model servers equivalent to the documented ones (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114 instead of the compose llm service); model-server startup itself not re-verified. Receipts were attested (signed by the box's key). The C2PA credential answered 503 as documented: the default image has no c2pa-python and no certificate. Host networking and port 8437 were used because other services held the default ports."},"known_limits":["Hosted check ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, signed receipts): console flow, public ledger page, text check, tamper buttons, Watch replay, 390 px layout, and the Build tab's Python example as written. The production API runs it once the branch is merged and deployed.","The claim check is a model's reading: on IteraTeR it caught 31 of 35 meaning-changed edits and also fired on many 'clarity' edits (most of which did change a fact).","The named editor's identity is what the caller sends; only an optional Ed25519 editor signature binds it to a key.","The C2PA credential uses a development certificate (untrusted issuer) and needs the provenance extra in the image.","Hosted retention is fixed at 7 days for open pieces and 30 days for published ledgers.","Under heavy shared load a whole piece took up to a minute."],"receipt_coverage":"full"},"cost_per_run_usd":0.002,"rehearsal_bundle":{"url":"/samples/editorial-ledger.zip","checks":9,"bytes":3627},"models":[{"name":"decosa-api editorial module (decosa_api/verticals/editorial)","role":"Ledger: edit metrics, sentence alignment, hash chain, sign-off rules, sealing and verification (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Writes the AI first draft from the sources, and lists the claims an edit added or removed","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"c2pa-python 0.37 (native c2pa-rs)","role":"Optional C2PA content credential on a DOCX copy of the published text","license":"MIT OR Apache-2.0","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/editorial-ledger","page":"/tools/media/editorial-ledger","json":"/use-cases/editorial-ledger.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"report-integrity","num":"32","name":"Report integrity","status":"live","industries":["public-sector","legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted additions found","value":"16/16","unit":null,"n":16,"split":"test","note":null},{"name":"Planted contradictions found (exact)","value":"16/16 (16/16)","unit":null,"n":16,"split":"test","note":null},{"name":"Planted omissions found, missing or partly (missing only)","value":"16/16 (13/16)","unit":null,"n":16,"split":"test","note":null},{"name":"Faithful sentences flagged (hard: unsupported or contradicted)","value":"0/64 (0)","unit":null,"n":64,"split":"test","note":null},{"name":"Key events flagged on faithful reports, missing or partly (missing only)","value":"9/68 (2/68)","unit":null,"n":68,"split":"test","note":null},{"name":"Plants found on ASR transcripts of the demo pair","value":"12/12","unit":null,"n":12,"split":"synthetic","note":"Two reports checked against MOSS-Transcribe-Diarize transcripts of synthetic audio, re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 2/19 faithful sentences flagged, one from a misheard name (first build: 12/12, 1/19)."}],"dataset":"12 invented incidents (5 police, 3 security, 2 EMS, 2 workplace), each with a timestamped transcript, a faithful report and a tampered report with six planted discrepancies (2 contradictions, 2 additions, 2 omissions). Dev 4 cases, test 8 cases run once after the prompts were frozen.","held_out":true,"caveats":["The building agent wrote the scenarios and the plants, so they are clear-cut; real reports are messier.","Synthetic incidents only: not measured on real incident reports against real body-worn-camera audio.","One run of the test split.","The check trusts the transcript: a mishearing the writer turns into a fact is marked as in the recording.","The first-draft writer's quality and the lite and best tiers are not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/report-integrity"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out test, 8 synthetic incidents: planted additions / contradictions / omissions found","value":"16/16 / 16/16 / 16/16","source":"decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route, one run; data written by the building agent, prompts frozen on a separate 4-case dev split"},{"metric":"faithful reports: sentences flagged / key events flagged as left out","value":"0/64 / 9/68 (2/68 as missing, 7 as partly)","source":"decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route"},{"metric":"the two demo incidents on real speech-recognition transcripts","value":"4/4 additions, 4/4 contradictions, 4/4 omissions found; faithful sentences flagged 2/19 (1 partial, 1 contradicted)","source":"decosa-api docs/evals/report-integrity.md (asr run), measured on our server 2026-09-26, gateway route: the traffic-stop and bar-fight reports checked against MOSS-Transcribe-Diarize transcripts of their synthetic audio, re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 1/19 flagged, partial)"},{"metric":"real reports against real body-worn-camera audio, reviewed by a lawyer","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":97000,"p95_ms":null,"runs":null,"receipts_per_run":21,"cost_per_run_usd":0.004},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"Verified on 2026-09-25: the image builds, the service starts, and both smoke tests in the assembly prompt pass end to end against a local model server equivalent to the documented one (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114, reached with network_mode host instead of the compose llm service); model-server startup itself not re-verified, and the diarizer path was not run. Check: 2 contradicted, 2 unsupported, refused search and the one-beer answer missing, 20 attested receipts, signature verified, 6 s. Draft: 12 sentences with two likely mishearings listed for the author, finalize 12 AI / 1 human, record verified, 13 s."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route) driven from the branch site in headless Chromium, including 390 px; the production API gets this vertical when the branch merges.","The check trusts the transcript. Speech-recognition errors pass through: in the demo the diarizer heard 'Camera on' as 'Cameron', and the draft named a bartender Cameron; the check marks it 'in the recording'. The writer now lists likely mishearings for the author to confirm.","'Not in the recording' is not 'false': what a camera cannot hear (smells, what someone saw) is flagged and needs the author's own account.","Measured on 12 synthetic incidents written by the building agent, with clear-cut plants; real reports are messier. Not yet measured on real body-worn-camera audio or with a lawyer reviewing.","Key-event coverage is noisier than the sentence check: 9 of 68 events on faithful reports were marked partly or missing, mostly detail a reader would not miss.","Hosted runs took 80-130 s while the shared gateway was busy (8-36 s when quiet).","Audio intake (/report/transcribe) is self-host only."],"receipt_coverage":"full"},"cost_per_run_usd":0.004,"rehearsal_bundle":{"url":"/samples/report-integrity.zip","checks":13,"bytes":4050},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Key events, first-draft writer, sentence judge (the grounding checker) and event coverage","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/report-integrity","page":"/legal/report-integrity","json":"/use-cases/report-integrity.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"legal-drafting-editor","num":"33","name":"Privileged drafting editor","status":"live","industries":["legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Clause detection vs CUAD labels, v1 mapping (like for like): precision / recall","value":"0.77 / 0.79","unit":null,"n":30,"split":"heldout","note":"30 CUAD test contracts, 8 rules, run once. v1 was 0.78 / 0.82: no gain like for like."},{"name":"Clause detection vs CUAD labels, v2 mapping: precision / recall","value":"0.82 / 0.85","unit":null,"n":30,"split":"heldout","note":"Part of the gain is the rule redefinition (new CONSEQ rule matches CUAD 'Cap On Liability'), not better reading."},{"name":"Planted deviations flagged: precision / recall","value":"0.98 / 1.00","unit":null,"n":102,"split":"heldout","note":"52 of 52 deviations flagged, 1 false flag; 17 contracts could be planted."},{"name":"Exact position of a planted clause (5 classes)","value":"0.98","unit":null,"n":102,"split":"heldout","note":"100 / 102"},{"name":"Inserted text not traceable to a precedent","value":"0 in 136","unit":"tracked changes","n":136,"split":"heldout","note":null},{"name":"LibreOffice Accept All matches our proposal","value":"46 / 47","unit":null,"n":47,"split":"heldout","note":"Mismatch was a deleted last paragraph; writer bug fixed after the run, test set not re-run."},{"name":"Model cost per contract (median, list price)","value":"US$0.0089","unit":null,"n":null,"split":"heldout","note":"16 calls, 22.9k prompt + 1.5k output tokens"}],"dataset":"CUAD v1 (510 SEC EDGAR contracts, expert clause labels, CC BY 4.0): 40 train-split dev contracts for tuning, 30 test-split contracts run once after the configuration was frozen. Playbook and precedents are synthetic; planted deviations from a clause-template bank.","held_out":true,"caveats":["The playbook and precedent library are synthetic, written for this demo; only clause presence and location come from CUAD.","Per-rule samples are small (3 to 6 labelled contracts per rule): warranty duration found in 0 of 3 and insurance in 2 of 5.","Only 17 of 30 test contracts had a buyer side, so only those could be planted.","Not measured: real firm playbooks, Microsoft Word itself, agreement with a lawyer's own redline, latency on a dedicated card.","One writer bug found by the test run was fixed afterwards; the test set was not re-run."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/legal-drafting-editor"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card, self-hosted","evidence":[{"metric":"Findings and redlines on CUAD","value":"not measured separately: the same weights and prompts as the standard tier; speed on a 5090 not measured","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Clause detection vs CUAD expert labels, 8 units × 30 held-out test contracts: precision / recall","value":"0.82 / 0.85 under the v2 mapping (v1: 0.81 / 0.79, re-scored); like for like with v1's mapping 0.77 / 0.79 (v1: 0.78 / 0.82), so no gain there","source":"docs/evals/legal-drafting-editor.md (v2). The v2 gain is the new CONSEQ rule, which finds the exclusions of indirect damages CUAD files under \"Cap On Liability\" (9 of 9); warranty duration 0 of 3 and insurance 2 of 5 labelled contracts found"},{"metric":"Planted deviations flagged (102 planted clauses, 17 CUAD test contracts): precision / recall","value":"0.98 / 1.00 (52 of 52 deviations flagged, 1 false flag); exact position (of 5) 0.98","source":"docs/evals/legal-drafting-editor.md; v1 was 0.89 / 0.95, exact 0.85, on 13 contracts, and 0.95 / 0.97, exact 0.92, once 6 items a harness bug had left out of the file are removed"},{"metric":"Contracts where the buyer side was named (needed for party-specific precedents)","value":"17 of 30 (v1: 13); both names among CUAD's labelled parties in 17 of 17","source":"docs/evals/legal-drafting-editor.md"},{"metric":"Redlines that are valid, reject-all = input, accept-all = proposal","value":"47 of 47","source":"docs/evals/legal-drafting-editor.md, every output file of the v2 held-out run"},{"metric":"Redlines LibreOffice opens and whose Accept All / Reject All match","value":"47 of 47 open, Reject All 47 of 47, Accept All 46 of 47; 284 of 284 comments imported","source":"docs/evals/legal-drafting-editor.md: the mismatch was a deleted last paragraph (LibreOffice and Word keep an empty one); fixed after the run and tested, not re-run on the test set"},{"metric":"Inserted text not traceable to a firm precedent","value":"0 in 136 tracked changes","source":"docs/evals/legal-drafting-editor.md; enforced in code: an untraced chunk sends the edit back to the precedent's sentence"},{"metric":"Targeted edits the checks sent back to the precedent's own sentence","value":"26 of 77","source":"docs/evals/legal-drafting-editor.md (v2 held-out run)"},{"metric":"Model cost per contract, eleven rules (median)","value":"US$0.0089 at list price ($0.30 / $1.50 per million tokens), 16 calls (v1: US$0.0065, 14 calls)","source":"docs/evals/legal-drafting-editor.md, 47 runs"},{"metric":"Levers tried on 40 dev contracts and left off","value":"Self-consistency (3 readings): 2.6x the cost, no detection gain. Topic gate: +0.05 precision, -0.06 recall. Embedding retrieval (bge-small, CPU): +0 to +1 point of labelled text shown. Log-probabilities: the gateway does not return them.","source":"docs/evals/legal-drafting-editor.md, lever study"}]},{"tier":"wanted","label":"Wanted · a larger second reader on your own hardware","evidence":[{"metric":"CUAD precision and recall on the same 8 rules","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":28000,"p95_ms":null,"runs":null,"receipts_per_run":20,"cost_per_run_usd":0.0066},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds, the service starts healthy, the Northwind sample passes end to end (9.7 s, 20 attested calls, 7 tracked changes, 0 untraced insertions, the file's checks pass), the signed record verifies and fails when one finding is changed, and the .docx export downloads. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["The playbook and precedent library in the demo are synthetic; a firm's own need to be loaded (as JSON) and were not tested.","Paragraphs with fields, hyperlinks, drawings or earlier tracked changes get a comment, not an in-place edit; inserted clauses are not numbered.","Checked against the OOXML rules and LibreOffice; not yet opened in Microsoft Word by us.","Position errors: 2% of planted clauses got the wrong one of five positions (v2 held-out run) and no planted deviation was missed; like-for-like clause detection against CUAD is 0.77 / 0.79, and warranty-duration and insurance clauses are the ones most often missed.","The client-side playbook needs a buyer side: contracts between equals (joint ventures, cooperation agreements) get findings but party-specific precedents stay in comments."],"receipt_coverage":"full"},"cost_per_run_usd":0.0066,"rehearsal_bundle":{"url":"/samples/legal-drafting-editor.zip","checks":12,"bytes":11768},"models":[{"name":"decosa-api drafting module (decosa_api/verticals/drafting)","role":"Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clause","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/legal-drafting-editor","page":"/legal/legal-drafting-editor","json":"/use-cases/legal-drafting-editor.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"foia-desk","num":"34","name":"Public-records desk","status":"live","industries":["public-sector","legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Personal-data spans redacted (PII recall), final pipeline","value":"59 of 60 (98.3%)","unit":null,"n":60,"split":"test","note":"Test run 5; the test sets were run five times with fixes in between."},{"name":"Planted spans redacted, final pipeline","value":"70 of 71 (98.6%)","unit":null,"n":71,"split":"test","note":"Pattern finders alone: 33 of 71 (46.5%)."},{"name":"Redaction precision / keep violations","value":"83.1% (69 of 83) / 3","unit":null,"n":83,"split":"test","note":"Most over-redactions are company officials' names and business emails."},{"name":"Responsiveness precision / recall","value":"96.2% / 100.0%","unit":null,"n":77,"split":"test","note":null},{"name":"Withheld spans recoverable from the release PDF","value":"0 of 95","unit":null,"n":95,"split":"test","note":"All readers; an earlier run (test run 2) failed this check."},{"name":"PII recall on a set new to the pipeline (bus-contract, run 3)","value":"5/8","unit":null,"n":8,"split":"heldout","note":"Responsiveness 7/8; the cleanest held-out number, before the medical-span fix."},{"name":"Cost at list price","value":"$0.77 per 1,000 records","unit":null,"n":77,"split":"test","note":null}],"dataset":"Synthetic agency records (3 dev sets, 4 held-out sets, 48 held-out records) labelled by Claude, plus 29 public FERC-released Enron emails labelled by hand for responsiveness.","held_out":true,"caveats":["The test sets were run four times and three design faults were fixed after seeing test results; the final numbers are not clean held-out numbers.","Labels were written by one AI reviewer (Claude), not a records officer.","Small sample: 48 synthetic held-out records and 29 public emails.","Synthetic records are cleaner than real ones: no OCR noise, attachments or long threads.","Medical details are the weakest kind (5/7); an officer still has to read each released record.","Only the hosted route was measured; the self-host route was smoke-tested."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/foia-desk"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card, self-hosted","evidence":[{"metric":"Held-out test sets","value":"not measured separately: the same weights and prompts as the standard tier; speed on a 5090 not measured","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Planted personal details redacted, 48 held-out synthetic records (60 labelled: names, phones, addresses, emails, SSN, dates of birth, licence and account numbers, medical details)","value":"59 of 60 (98.3%); the pattern finders alone would catch 46.5% of all planted spans","source":"docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline); labels written by Claude (an AI agent), not a records officer"},{"metric":"What it missed","value":"The misses are listed per record in the eval write-up; medical details are the weakest kind.","source":"docs/evals/foia-desk.md"},{"metric":"Responsiveness, 77 held-out records (48 synthetic, 29 public FERC-released Enron emails): precision / recall","value":"96.2% / 100.0%","source":"docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)"},{"metric":"Exemption label on redacted spans / records with agency counsel withheld in full with the right exemption","value":"70 of 70 / 6 of 6 (federal (b)(5), (b)(6), (b)(4); CPRA § 7927.700, § 7927.705, § 7922.000)","source":"docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)"},{"metric":"Redaction precision (redactions that overlap a labelled span) / text that had to be released but was boxed","value":"83.1% / 3 spans: mostly company officials' names and emails, boxed everywhere by the privacy-protective consistency step and sent to review","source":"docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)"},{"metric":"Release check: withheld spans recoverable from the PDF (built-in reader, pdfplumber, filing-preflight black-box check)","value":"0 of 95 on the held-out releases; 0 characters under any box","source":"docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)"},{"metric":"How the numbers moved as design faults were fixed (the test sets were run five times; no threshold tuned on them)","value":"personal details redacted 81.0% → 78.8% → 91.7% → 96.7% → 98.3%; responsiveness recall 91.4% → 86.0% → 98.0% → 100.0% → 100.0%","source":"docs/evals/foia-desk.md, History; two fresh held-out sets were written before runs 2 and 3"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":40000,"p95_ms":null,"runs":null,"receipts_per_run":41,"cost_per_run_usd":0.012},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs and both exemption lists, scoping works, the coastal-permits sample passes end to end (17 s, 41 calls, every disposition as expected, release check ok on 28 spans), the signed record verifies and fails at the changed entry, the PDF export carries no withheld text, one officer decision rebuilds the release and the ledger verifies. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Labels are one AI reviewer's (Claude's), not a records officer's; the held-out sets are small and synthetic.","Medical details are the weakest kind: both held-out misses were medical.","It over-redacts company officials' names that the model reads as private, and agency staff named only in a record's text (not in its mail headers); they go to review or to the officer.","Runs are not identical: the same records can come back with a few redactions more or fewer from one run to the next (temperature 0 on a shared server is not bit-for-bit repeatable). The pattern finders, the date filter and the last contact check are code and do not vary; the officer decides every proposal.","The release is re-typeset from text: no original layout, no PDF, scan or email-archive intake, no audio or video.","The hosted route slows when the shared gateway is loaded."],"receipt_coverage":"full"},"cost_per_run_usd":0.012,"rehearsal_bundle":{"url":"/samples/foia-desk.zip","checks":11,"bytes":3746},"models":[{"name":"decosa-api foia module (decosa_api/verticals/foia)","role":"Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's calls","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/foia-desk","page":"/legal/foia-desk","json":"/use-cases/foia-desk.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"oral-assessment","num":"35","name":"Structured oral assessment","status":"live","industries":["education","hr-recruiting"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Draft level equals the label (exact)","value":"90.8%","unit":null,"n":120,"split":"test","note":"Within one level: 100%. 120 criteria come from 60 distinct answer-criterion pairs."},{"name":"Quadratic weighted kappa vs labels","value":"0.971","unit":null,"n":120,"split":"test","note":null},{"name":"Misconception or mixed answers scored exactly","value":"70.8%","unit":null,"n":null,"split":"test","note":"Strong 100%, partial 91.7%, weak 91.7%."},{"name":"Fluency penalty: plain non-native English vs fluent strong answers","value":"none measured (24 of 24 each)","unit":null,"n":24,"split":"test","note":"Written text, not real accented speech."},{"name":"Test-retest, same level","value":"118 of 120 (98.3%)","unit":null,"n":120,"split":"test","note":null},{"name":"Real recognition errors found by the disagreement marks: precision / recall","value":"65% / 74%","unit":null,"n":36,"split":"synthetic","note":"27 errors on synthetic TTS voices, clean and in noise. Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M, American and British English only) and re-run; the first build (macOS voices, six English accents): 53% / 81%, 36 errors."}],"dataset":"20 synthetic transcripts (10 per rubric: intro-statistics viva and customer-support interview), 6 criteria each, 120 scored criteria per run; labels written with the answers before any model run. Prompts developed on the four demo scripts only; the test set was not used for tuning.","held_out":true,"caveats":["Labels are one AI author's (the building agent), who also wrote the answers to hit the levels: agreement here is an upper bound on real vivas.","Synthetic only; no examiner labels and no examiner-examiner agreement to compare with.","Review flags did not predict the model's disagreements (0 of 11 flagged); the examiner must decide every criterion.","Speech measured only on synthetic voices; real accented, fast or overlapping speech will have many more recognition errors.","Only the two built-in rubrics were tested."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/oral-assessment"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card, captions only","evidence":[{"metric":"scoring agreement from captions with guessed roles","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, two recognisers","evidence":[{"metric":"draft level vs labels, 120 held-out criteria: exact / within one / QWK","value":"90.8% / 100% / 0.971","source":"decosa-api docs/evals/oral-assessment.md, 2026-09-25; synthetic answers, labels by Claude (an AI agent), not examiners"},{"metric":"fluency penalty: strong answers in plain, non-native English","value":"none measured (24 of 24 same level)","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"evidence: levels above 0 citing the right answer; grounding supported","value":"100%; 99.2%","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"test-retest same level","value":"98.3%","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"10% of words mis-recognised: levels changed; lowered scores flagged when marked","value":"7.5%; 11 of 12","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"real recognition errors found by the disagreement marks","value":"74% recall, 65% precision (27 errors, synthetic voices)","source":"decosa-api docs/evals/oral-assessment.md"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash scores and checks","evidence":[{"metric":"scoring agreement","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":17600,"p95_ms":null,"runs":null,"receipts_per_run":42,"cost_per_run_usd":0.007},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's api service (named volume) against the running local model servers, then torn down","notes":"The step 6 smoke passed as written: six scores with cited lines, the leading question at line 3 flagged, one override signed, the bundle verified, and after editing the override's reason verification failed at that decision. The audio replay also passed (speaker lines, signed draft of 87 entries). Model-server startup itself not re-verified (no new GPU load)."},"known_limits":["Scores are drafts. On 120 held-out synthetic criteria they matched labels written by an AI agent (Claude) 90.8% of the time and were never more than one level off; real answers and real examiners are not measured yet.","The review flags did not catch the model's disagreements (0 of 11), and on the hosted route the stated confidence is nearly always 0.95; the examiner has to decide every criterion.","Speech was measured on synthetic TTS voices only. Words the two recognisers disagree on are marked (81% of real errors found in the eval), but accented or overlapping human speech is untested.","Rubrics are JSON: two built in, custom ones through the API; no rubric editor in the console yet.","No LMS or ATS export yet; keep the bundle JSON with the grade."],"receipt_coverage":"full"},"cost_per_run_usd":0.007,"rehearsal_bundle":{"url":"/samples/oral-assessment.zip","checks":11,"bytes":4137},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Live captions (streaming, no speakers): the examiner prompt reads these","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"After the session: examiner and candidate lines, each with a receipt over its audio","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Examiner prompt, speaker roles, one score per criterion, grounding check of the evidence","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/oral-assessment","page":"/tools/operations/oral-assessment","json":"/use-cases/oral-assessment.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"patent-claim-support","num":"36","name":"Patent claim-support checker","status":"live","industries":["legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Supported-or-not call agrees with the labels","value":"93.4%","unit":null,"n":166,"split":"test","note":"100% on the 146 elements labelled as clear calls. Three runs: 93.6%, 91.0%, 93.4%. Dev: 97.8%."},{"name":"Supported elements where a cited paragraph is one the labels list","value":"100%","unit":null,"n":150,"split":"test","note":"Share of cited paragraphs the labels list: 93.8%"},{"name":"Partly supported elements flagged (recall) / flags that match a label (precision)","value":"5 of 13 / 5 of 8","unit":null,"n":166,"split":"test","note":"Partial support is what it misses; all misses and false flags are judgment calls."},{"name":"Planted support removal flagged","value":"5 of 6","unit":null,"n":6,"split":"test","note":"Over two runs 10 of 12; both misses cite a paragraph that still describes the element."},{"name":"Antecedent-basis flags that are genuine, random granted claims","value":"55% (36 of 66)","unit":null,"n":66,"split":"heldout","note":"Round 4: 55 patents, 891 claims, drawn after the last rule change. About 1.2 flags per patent."},{"name":"Cost per application (list price)","value":"$0.053","unit":null,"n":null,"split":"test","note":"166 calls over 6 patents, about 6,000 prompt tokens per element"}],"dataset":"Eight granted US patents (public domain): 2 dev, 6 test (212 claim elements; 166 test), labelled per element by two separate Claude agents. Antecedent set: 213 randomly drawn granted patents in four rounds (3,478 claims); rounds 1-3 used to improve the rules, round 4 held out.","held_out":true,"caveats":["Labels are by AI agents, not a registered practitioner; no examiner reasons for allowance to check against.","Granted claims were examined, so most elements are expected to be supported.","Partial support is often missed (5 of 13): treat \"supported\" as \"here is where to look\", not \"no 112(a) problem\".","Antecedent recall is measured only on planted errors; precision may differ in other art units.","Not measured: long specifications over 48,000 characters, the 32 GB card tier."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/patent-claim-support"},"quality_evidence":[{"tier":"lite","label":"Lite · code checks only, any CPU","evidence":[{"metric":"Antecedent-basis flags that are genuine, on randomly drawn granted US claims (held out: rules frozen before the set was drawn)","value":"54.5% (36 of 66 flags; 55 patents, 891 claims)","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25; each flag judged by a separate Claude agent with a written rubric, not a practitioner"},{"metric":"Planted defects found by code in 6 test patents: antecedent errors / broken claim references / renamed terms","value":"11 of 11 / 12 of 12 / 6 of 6","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Elements where the model's supported-or-not call matches the labels, 6 held-out granted patents","value":"93.4% of 166 (100.0% of the 146 labelled as clear calls)","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25; labels written by Claude (an AI agent) reading each specification, not by a practitioner"},{"metric":"Supported elements where a cited paragraph is one the labels list / share of cited paragraphs that are","value":"100.0% / 93.8% (150 elements)","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25"},{"metric":"Elements the labels call only partly supported that the model also flags (recall) / flags that match a label (precision)","value":"5 of 13 / 5 of 8","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25; most misses are limitations the labels mark as judgment calls, such as a feature described only in a different embodiment"},{"metric":"Planted support removal: all paragraphs describing one element deleted, element flagged","value":"5 of 6 (the miss: a paragraph the removal left in still describes the element, checked by reading it; an earlier run missed a different patent the same way)","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25"},{"metric":"Dev patents (2, used to write the prompt): agreement / flag recall","value":"97.8% / 3 of 3","source":"decosa-api docs/evals/patent-claim-support.md, 2026-09-25"}]},{"tier":"wanted","label":"Wanted · a GLM-5.3-Flash judge on your own hardware","evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":7800,"p95_ms":null,"runs":null,"receipts_per_run":13,"cost_per_run_usd":0.0204},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds, the service starts healthy, the planted sample runs end to end on the direct route (13 attested calls, 3.4 s): claims 2 and 5 have no support, the antecedent, reference and term issues are found, the signed record verifies and fails when one strength is changed, the CSV export works, and no claim text reaches the logs. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Labels are Claude's (an AI agent), not a registered practitioner's; antecedent flags were judged by separate Claude agents with a written rubric.","The support map finds supporting paragraphs well but flags only 5 of the 13 held-out elements the labels call partly supported; most misses are judgment calls such as a feature described only in another embodiment.","About half of the antecedent-basis flags on randomly drawn granted claims are genuine slips (36 of 66 held out); the rest are implicit or inherent references a practitioner would accept.","Run-to-run variation on the hosted gateway: the same element can come back partly supported in one run and unsupported in the next; both are flagged.","Text only (no DOCX, PDF or figures); specifications over 48,000 characters are read as the best-matching paragraphs, which can miss support far from the element's words."],"receipt_coverage":"full"},"cost_per_run_usd":0.0204,"rehearsal_bundle":{"url":"/samples/patent-claim-support.zip","checks":13,"bytes":15456},"models":[{"name":"decosa-api patent module (decosa_api/verticals/patent)","role":"Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: reads each claim element against the specification and names the supporting paragraphs","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/patent-claim-support","page":"/legal/patent-claim-support","json":"/use-cases/patent-claim-support.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"music-gen-cleared","num":"37","name":"Rights-cleared music generation","status":"live","industries":["music","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Prompt guard precision (refuse)","value":"100% (48/48)","unit":null,"n":79,"split":"test","note":"Dev: 96.3% (26/27)"},{"name":"Prompt guard recall (refuse)","value":"98.0% (48/49)","unit":null,"n":79,"split":"test","note":"Overt 26/26, sneaky 22/23; code rules alone refuse 55% (27 of 49)."},{"name":"Safe prompts refused","value":"0 of 30","unit":null,"n":30,"split":"test","note":"Dev: 1 of 15 (a public-domain composer refused)."},{"name":"Planted near-copies flagged at the flag threshold","value":"80.8%","unit":null,"n":240,"split":"test","note":"90.4% at the review threshold. About one planted copy in five slips under the flag."},{"name":"False flags on unrelated tracks by new artists","value":"2.3% (2)","unit":null,"n":88,"split":"test","note":"Held-out pool tracks: 1.9% (1 of 53); fresh renders 0 of 9."}],"dataset":"Prompt guard: 120 hand-written synthetic prompts (45 safe, 40 overt, 35 sneaky), 41 dev / 79 test, written before the guard ran. Similarity: 240 commercially licensed Free Music Archive tracks as the catalogue; 64 planted near-copies x 6 edits (24 dev / 40 test sources), plus unrelated and held-out negatives; thresholds set on dev only, applied once to test.","held_out":true,"caveats":["The guard prompts were written by hand for this eval by the builder; synthetic.","The similarity check is a triage signal for near-copies of the references it holds, not a clearance: 240 Creative Commons tracks are not the recordings anyone is likely to copy.","Lyrics are not compared.","Music quality is not measured.","Guard latency (median 30.3 s) was measured on a saturated shared gateway."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/music-gen-cleared"},"quality_evidence":[{"tier":"lite","label":"Lite · MIT engine, about 15 GB of GPU","evidence":[{"metric":"Prompt guard on held-out prompts (79: 30 safe, 26 overt, 23 sneaky)","value":"precision 100%, recall 98% (48 of 49), 0 of 30 safe prompts refused","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25; prompts hand-written for the eval"},{"metric":"Planted near-copies flagged (held-out test, 240 clips)","value":"80.8% at 2.3% false alarms on unrelated tracks (2 of 88)","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25; threshold set on the dev split"},{"metric":"Music quality","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, MiniMax-Music3 on a 96 GB card","evidence":[{"metric":"Prompt guard on held-out prompts (79)","value":"precision 100%, recall 98%; code rules alone: recall 55%","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25"},{"metric":"Planted near-copies flagged, by change (held-out)","value":"pitch +2: 75%, pitch -1: 85%, tempo 0.92: 83%, tempo 1.08: 85%, pitch +1 and tempo 1.05: 85%, 12 s excerpt pitched -2: 73%","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25"},{"metric":"Music quality","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":42700,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.00019},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image plus the documented FFmpeg and c2pa-python layer and of services/music_embed, compose with named volumes, pointed at the already-running local vLLM (direct route) and ComfyUI; then torn down.","notes":"The guard refused and stripped as documented, a 30 s MiniMax-Music3 render finished in 21.4 s with a similarity verdict and a certificate that verified (and failed when edited), a planted near-copy scored 11.6 against the catalogue, and no prompt text reached the logs. Found on the way: the api image has no c2pa-python, so the first render had no C2PA credential; with the layer and the dev certificate, stamping inside the container produced a credential that validates."},"known_limits":["The similarity check knows only 240 Creative Commons tracks and the files you compare; it flagged 81% of planted near-copies at 2% false alarms, and a clear result says nothing about commercial catalogues.","The guard reads text only: 98% recall on held-out prompts, and it cannot hear a melody someone describes or hums.","Renders share one GPU with the live demos: 3 per demo session, 3 per API key per day.","C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators.","No copyright is claimed in the output; the certificate records checks and licence terms, not ownership."],"receipt_coverage":"partial"},"cost_per_run_usd":0.00019,"rehearsal_bundle":{"url":"/samples/music-gen-cleared.zip","checks":12,"bytes":517844},"models":[{"name":"decosa-api music module (decosa_api/verticals/music)","role":"Prompt guard, code layer: gazetteer of well-known names, look-alike folding, cover/cloning/type-beat patterns, a genre and era allow-list; the certificate and refusal log (CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Prompt guard, typed judgment: names the category (artist, work, voice, label or none) with a calibrated probability; one call per brief","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MiniMax-Music3","role":"Music generation: songs with vocals and lyrics (default engine)","license":"MiniMax-Music3 Community License","hf_repo":"MiniMaxAI/MiniMax-Music3"},{"name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","role":"Music generation: fast drafts and instrumentals (MIT engine)","license":"MIT","hf_repo":"ACE-Step/Ace-Step1.5"},{"name":"LAION CLAP larger_clap_music + librosa chroma (services/music_embed)","role":"Similarity check (CPU): melody by chroma alignment across keys and tempos, sound by CLAP window embeddings, both normalised per reference","license":"Apache-2.0","hf_repo":"laion/larger_clap_music"}],"licence":"community","links":{"metrics":"/metrics/music-gen-cleared","page":"/apps/music-gen-cleared","json":"/use-cases/music-gen-cleared.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"sample-clearance","num":"38","name":"Sample and lyric clearance pre-check","status":"live","industries":["music","legal"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Audio plants found, medium and high confidence","value":"62 of 108 (57%)","unit":null,"n":108,"split":"test","note":null},{"name":"Audio precision, medium and high confidence","value":"0.87 (9 false items)","unit":null,"n":null,"split":"test","note":null},{"name":"Audio plants found, high confidence only (precision)","value":"35 of 108 (32%), precision 1.00","unit":null,"n":108,"split":"test","note":null},{"name":"Clean tracks with any flag","value":"3 of 30","unit":null,"n":30,"split":"test","note":null},{"name":"Lyric plants found with the model's labels: exact / near / heavy paraphrase","value":"8/8, 12/12, 1/8","unit":null,"n":28,"split":"test","note":"1 false item; 0 of 15 clean songs flagged."},{"name":"Chromaprint baseline, samples found","value":"4 of 84 (616 chance matches)","unit":null,"n":84,"split":"test","note":null}],"dataset":"Audio: 35 CC BY 4.0 Kevin MacLeod reference recordings and 73 host recordings by the same artist, with synthetic plants (mixed slices under pitch, tempo, filter, loop and varispeed transforms, plus re-played melody interpolations) and clean excerpts; dev and test use different hosts and seeds. Lyrics: 156 US public-domain songs (5,435 lines) planted into 60 model-written songs, 30 dev and 30 test.","held_out":true,"caveats":["Samples are mixed synthetically; real productions add compression, reverb and other layers. No real-world recall is claimed.","Thresholds were tuned on dev; the test split was scored twice, before and after an interpolation rule changed on dev (first run: 62 of 108, precision 0.89).","One artist's music throughout; a catalog of many artists may behave differently.","Only references in the catalog can be found; the hosted catalog is a 41-reference demo.","Heavy lyric paraphrases are mostly out of reach; the lyric set is English and pre-1929."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/sample-clearance"},"quality_evidence":[{"tier":"lite","label":"Lite · code only, any CPU","evidence":[{"metric":"Planted samples and melodies found / precision / clean tracks flagged (test split, 108 plants in CC BY music, 30 clean tracks)","value":"62 of 108 (57%) / 0.87 / 3 of 30","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25 (thresholds set on the dev split; test scored twice, before and after one dev change, both reported)"},{"metric":"High-confidence items only: found / precision / clean tracks flagged","value":"35 of 108 / 1.00 / 0 of 30","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25"},{"metric":"By transform (test): pitch ±1-2 semitones / tempo ±5-10% / low-pass / high-pass / 2 s loops / interpolations","value":"14 of 24 / 16 of 24 / 5 of 6 / 6 of 6 / 2 of 6 / 12 of 24","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25"},{"metric":"Lyric lines, code only (test): exact / one or two words changed / heavy paraphrase / clean songs flagged","value":"8 of 8 / 12 of 12 / 0 of 8 / 0 of 15","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25 (with embeddings; without them, trigrams alone)"}]},{"tier":"standard","label":"Standard · adds lyric labels by Qwen3.8-27B (hosted demo)","evidence":[{"metric":"Lyric lines with the model's labels (test): exact / one or two words changed / heavy paraphrase / false items / clean songs flagged","value":"8 of 8 / 12 of 12 / 1 of 8 / 1 / 0 of 15","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25; 12 model calls, 515 generated tokens for 30 songs"},{"metric":"Audio (same matcher as Lite)","value":"62 of 108 found, precision 0.87","source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25"}]}],"benchmark":{"title":"How well does it find planted borrowings?","intro":"Slices of 35 CC BY recordings were pitched, stretched, filtered or looped and mixed 3-12 dB under 73 other tracks, melodies were re-played on other instruments, and public-domain lyric lines were copied or re-worded into new songs. Thresholds were set on a dev split; these are the test split.","rows":[{"label":"Samples and melodies found (108 plants)","value":"57% (62)","detail":"precision 0.87; 3 of 30 clean tracks got a medium-confidence flag"},{"label":"High-confidence items","value":"35 found, all correct","detail":"0 of 30 clean tracks flagged"},{"label":"Samples mixed at -3 / -6 / -9 / -12 dB","value":"71% / 57% / 57% / 43%","detail":null},{"label":"Lyric lines: exact / small changes / heavy paraphrase","value":"8/8 · 12/12 · 1/8","detail":"no clean song flagged"},{"label":"Chromaprint alone, same samples","value":"4 of 84","detail":"and 616 chance matches: built for whole recordings"}],"points":[{"heading":"Where it fails","text":"Quiet slices and short loops, where too few of the reference's spectral peaks survive the mix, and pitch-plus-tempo changes that fall between the search grid points. Three of the nine false items on test were tracks from one series by the same composer that may share material."},{"heading":"What this does not show","text":"Synthetic mixes of one composer's CC BY music are not real productions with compression and effects, and the catalog is 41 references. No real-world recall is claimed."}],"source":"decosa-api docs/evals/sample-clearance.md, 2026-09-25"},"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":6732,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.000112},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone into a clean directory, docker build, the api service with a named volume, embedding files and the demo catalog fetched inside the container, pointed at the running local vLLM (Qwen3.8-27B) over host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds with ffmpeg and the clearance extra, the demo catalog builds in 81 s (download included), the planted sample returns the same five items and the declared-not-found entry as the hosted run in 5.4 s on the direct route (one attested call), the clean sample returns nothing, the signed record verifies and fails when one action is changed, and no lyric or title text reaches the logs. The first attempt found a bug (the demo track list was not in the image), fixed in the branch. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Recall is moderate: 57% of planted borrowings on the test split; quiet samples and short loops are missed most.","It only finds what is in the catalog it is given; the hosted demo catalog is 41 references.","The eval mixes are synthetic, from one composer's CC BY music; real productions may behave differently.","Lyric matching is English and needs a lyric set: the demo uses public-domain songs published before 1929."],"receipt_coverage":"partial"},"cost_per_run_usd":0.000112,"rehearsal_bundle":{"url":"/samples/sample-clearance.zip","checks":10,"bytes":664923},"models":[{"name":"decosa-api clearance module (decosa_api/verticals/clearance)","role":"Audio matcher: log-frequency landmarks with pitch and tempo search, peak-by-peak verification, melody shingles for composition references; lyric spans; declared vs detected; the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"all-MiniLM-L6-v2 (ONNX)","role":"Lyric-line embeddings: re-worded lines that share few characters with the original","license":"Apache-2.0","hf_repo":"sentence-transformers/all-MiniLM-L6-v2"},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: labels each near-duplicate lyric line lift, variant, stock phrase or different","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/sample-clearance","page":"/tools/media/sample-clearance","json":"/use-cases/sample-clearance.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"split-sheet-check","num":"39","name":"Split-sheet and metadata checker","status":"live","industries":["music"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted errors found","value":"43 of 43","unit":null,"n":43,"split":"test","note":null},{"name":"Precision, all catalogs","value":"64.8%","unit":null,"n":null,"split":"test","note":"Knock-on issues counted separately (9)."},{"name":"Precision, typed split sheets only","value":"19 of 19","unit":null,"n":19,"split":"test","note":"15 catalogs; no false alarm."},{"name":"Precision, catalogs with a scanned split sheet","value":"27 of 52","unit":null,"n":52,"split":"test","note":"All 25 false alarms trace to OCR."},{"name":"IPI extracted correctly from handwriting-style scans via OCR","value":"33.9%","unit":null,"n":59,"split":"test","note":"100% on typed sheets (162 writers)."},{"name":"Share sums and identifiers recomputed independently, disagreements","value":"0 on 151 share sums and 374 identifiers","unit":null,"n":null,"split":"test","note":null}],"dataset":"Synthetic, fictional catalogs from our own generator (2-3 song releases with split sheets in five styles including handwriting-font scans, DDEX ERN or distributor CSV, society registration export, contract excerpts), 0-3 planted errors each. Dev 8 catalogs; test 30 catalogs with 43 planted errors, run once after dev was frozen.","held_out":true,"caveats":["Synthetic data from our own generator; the generator and the checker were written by the same agent, so the structured formats are friendly to the parser.","Handwriting fonts are not handwriting; real handwritten sheets will do worse through Tesseract.","Scanned split sheets bring many OCR false alarms (25 of 52 issues).","One run; output at temperature 0 on a shared gateway can still vary slightly."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/split-sheet-check"},"quality_evidence":[{"tier":"lite","label":"Lite · code only, any CPU","evidence":[{"metric":"Identifiers and share sums in the reports rechecked by an independent implementation (disagreements)","value":"0 of 374 identifiers, 0 of 151 sums","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Planted identifier errors found (ISWC, IPI, UPC check digits; malformed ISRC)","value":"ISWC check digit 3/3; IPI check digits 4/4; UPC check digit 5/5; malformed ISRC 6/6","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Planted errors found, 30 held-out synthetic catalogs (43 planted)","value":"43 of 43 (100.0%): IPI check digits 4/4; malformed ISRC 6/6; ISWC check digit 3/3; missing co-writer 6/6; missing publisher 1/1; society mismatch 5/5; composer/lyricist swap 8/8; 105% split 5/5; UPC check digit 5/5","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Issues that are a planted error, not a false alarm (knock-on effects of a planted error excluded)","value":"64.8%: id_invalid 18/36; missing_writer 6/9; pro_mismatch 5/6; publisher_mismatch 0/2; publisher_missing 2/2; role_mismatch 8/8; share_total 7/7; share_unreadable 0/1","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Issues that are a planted error, split by whether the catalog has a scanned split sheet","value":"typed-only catalogs (15): 19 of 19, no false alarm; catalogs with a scan (15): 27 of 52, every false alarm traces to an OCR misread (22 of the 25 are warnings marked as OCR)","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Unknown role words (\"beat + chords\", \"wrote the lyrics\") resolved by the typed role question","value":"10 of 10 correct","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"False alarms on the 7 clean catalogs","value":"3","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Split-sheet values read correctly, typed sheets (162 writers)","value":"name 100.0%, share 100.0%, role 93.8%, ipi 100.0%, pro 100.0%, publisher 100.0%","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out"},{"metric":"Split-sheet values read correctly, handwriting-style scans through OCR (59 writers)","value":"name 98.3%, share 96.6%, role 98.3%, ipi 33.9%, pro 96.6%, publisher 94.9%","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), test seed held out; Tesseract 5 on synthetic handwriting fonts, not real handwriting"},{"metric":"Dev catalogs (8, used to fix the prompt and the OCR handling): planted found / precision","value":"10 of 10 / 83.3%","source":"decosa-api docs/evals/split-sheet-check.md, 2026-09-25; synthetic catalogs from scripts/splits_data.py (fictional), dev seed"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":6728,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.00303},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile (with Tesseract), the api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"The image builds with OCR on; the planted sample runs end to end on the direct route in 4.5 s (5 attested calls, the same 9 issues as hosted) and the scanned sample in 5.2 s (OCR in the container); the signed record verifies; no names or titles reach the logs. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Measured on synthetic, fictional catalogs only; real exports from each society and distributor have their own layouts.","Handwriting-style scans read through Tesseract lose identifiers (33.9% of IPIs read correctly); they are marked as OCR and not trusted.","A valid check digit says an identifier is well formed, not that it is registered to the person named.","One release at a time (12 documents hosted); no bulk catalog mode and no CWR filing.","Latency depends on the shared gateway: 30-100 s per release was measured while it was saturated."],"receipt_coverage":"full"},"cost_per_run_usd":0.00303,"rehearsal_bundle":{"url":"/samples/split-sheet-check.zip","checks":12,"bytes":127032},"models":[{"name":"decosa-api splits module (decosa_api/verticals/splits) with Tesseract OCR","role":"Checker: DDEX and CSV parsing, OCR, identifier check digits, exact share arithmetic, grounding of every model-read value, reconciliation, the proposed sheet and signed records (no model; CPU)","license":"AGPL-3.0-or-later (decosa-api; Tesseract 5)","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: reads split sheets and contract excerpts (values copied as written, with line numbers) and answers typed same-writer and role questions","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/split-sheet-check","page":"/tools/media/split-sheet-check","json":"/use-cases/split-sheet-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"paper-claim-check","num":"43","name":"Citation and claim checker for papers","status":"live","industries":["science-research"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted wrong-paper citations flagged","value":"12 of 12","unit":null,"n":12,"split":"test","note":"95% CI 76-100%"},{"name":"Planted overstatements flagged","value":"12 of 12","unit":null,"n":12,"split":"test","note":"All came back contradicted, with the differing passage quoted. 95% CI 76-100%."},{"name":"False alarms on correct citations with open full text","value":"12 of 49 (24.5%)","unit":null,"n":49,"split":"test","note":"95% CI 14.6-38.1%: about one flag in four is noise."},{"name":"Planted retracted citations / year errors / DOI swaps flagged","value":"6 of 6 / 6 of 6 / 6 of 6","unit":null,"n":18,"split":"test","note":"Year and DOI results are after fixes made following the test run (before: 5 of 6 and 4 of 6)."},{"name":"Retraction flag on known retracted works, with DOI / without","value":"68 of 68 / 68 of 68","unit":null,"n":68,"split":"heldout","note":"Drawn from Crossref's own notices: tests the pipeline, not Retraction Watch coverage."},{"name":"Retraction flag on random control articles, with DOI / without","value":"0 of 60 / 0 of 60","unit":null,"n":60,"split":"heldout","note":null}],"dataset":"Eight CC BY PubMed Central papers: 2 dev papers used to write the prompt, 6 test papers fixed before any test run (471 citation pairs), with planted errors; 60 unplanted citation pairs labelled blind; a retraction set of 68 retracted works and 60 random controls from Crossref/OpenAlex metadata.","held_out":true,"caveats":["False-alarm labels are an AI agent's, not a domain expert's, on only 49 correct citations: a wide interval.","Overstatement plants were hand-written by the builder.","Only open-access cited papers can be checked (44% of test pairs had full text); figures and tables are not read.","Reference-check fixes were made after the test run and re-run without the model; the claim verdicts are from the first run.","Run-to-run variation: the same sentence can be flagged in one run and supported in the next."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/paper-claim-check"},"quality_evidence":[{"tier":"lite","label":"Lite · references only, any CPU","evidence":[{"metric":"Retraction flag on known retracted works (60 sampled from Crossref's notices + 8 well-known), cited with DOI / without a DOI","value":"68 of 68 / 68 of 68","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the sampled set comes from the same Crossref data, so this tests the pipeline, not Retraction Watch's coverage"},{"metric":"Retraction flag on 60 random journal articles (controls), with DOI / without","value":"0 of 60 / 0 of 60","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25"},{"metric":"Planted retracted citations / wrong years / wrong DOIs found in 6 test papers","value":"6 of 6 / 6 of 6 / 6 of 6","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the no-DOI retraction and the DOI plants after fixes made on the test run (disclosed there)"},{"metric":"Unplanted reference flags on the same 6 real papers (431 references): genuine / arguable / wrong","value":"3 / 3 / 0 errors and warnings (two wrong DOIs in a published paper, one malformed author list; one ambiguous supplement DOI, two software records whose year differs), after the fixes","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25; judged by Claude (an AI agent)"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Planted wrong-paper citations flagged not supported or contradicted, 6 held-out CC BY papers","value":"12 of 12","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25"},{"metric":"Planted overstatements flagged (numbers inflated, association made causal, population or design changed)","value":"12 of 12, each with the differing passage quoted","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25; plants written by Claude (an AI agent)"},{"metric":"False alarms on correct citations with open full text (blind labels)","value":"12 of 49 flagged (24.5%, 95% CI 14.6-38.1%)","source":"decosa-api docs/evals/paper-claim-check.md, 2026-09-25; labels written by Claude (an AI agent), not domain experts, before the verdicts were seen"}]},{"tier":"wanted","label":"Wanted · GLM-5.3-Flash for the claim check","evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":9574,"p95_ms":null,"runs":null,"receipts_per_run":4,"cost_per_run_usd":0.0028},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume plus the GROBID container, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.","notes":"Verified on 2026-09-25: the image builds, the service starts healthy with GROBID parsing, the smoke test passes (10.5 s, 4 attested calls), the planted sample flags reference 43 retracted, the overstated [42] sentence and reference 1's year, the signed report verifies and fails when one verdict is changed, and no manuscript text reaches the logs. In one of three runs the swapped [36] citation came back supported."},"known_limits":["About one flag in four on correct background citations is noise (12 of 49 on the test papers); each flag quotes the passage, so a reader can dismiss it quickly.","Only open-access cited papers are read; 56% of test citations had only an abstract or nothing, and those are marked unavailable.","Labels and plants are Claude's (an AI agent), not a domain expert's, on six biomedical papers.","Run-to-run variation: the same citation can be flagged in one run and supported in the next.","PDF text loses superscript citation numbers; figures and tables are not read."],"receipt_coverage":"full"},"cost_per_run_usd":0.0028,"rehearsal_bundle":{"url":"/samples/paper-claim-check.zip","checks":10,"bytes":2306},"models":[{"name":"decosa-api papercheck module (decosa_api/verticals/papercheck)","role":"Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"GROBID 0.8.2 (CRF models)","role":"Reference parser (optional): splits each reference into authors, title, year and DOI","license":"Apache-2.0","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: reads each citing sentence against the cited paper's passages, then reviews its own flags","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/paper-claim-check","page":"/tools/life-sciences/paper-claim-check","json":"/use-cases/paper-claim-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"signed-lab-notebook","num":"44","name":"Signed lab notebook","status":"live","industries":["science-research","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Genuine synthetic exports that verify","value":"300 / 300","unit":null,"n":300,"split":"synthetic","note":"Median check 4.4 ms in Python."},{"name":"Outsider alterations caught (editing the export without keys)","value":"1,400 / 1,400","unit":null,"n":1400,"split":"synthetic","note":null},{"name":"Insider alterations caught, personal keys + TSA, export alone","value":"515 of 520","unit":null,"n":520,"split":"synthetic","note":"All caught when checked against an earlier export (1,400 / 1,400 across configurations)."},{"name":"Insider edits caught, account signatures + TSA, export alone","value":"29/200","unit":null,"n":200,"split":"synthetic","note":"Deletes 2/40, reorders 14/40, backdates 40/80, forged signatures 40/80."},{"name":"Real recorded export, insider alterations caught on the export alone","value":"10 of 12","unit":null,"n":12,"split":"synthetic","note":"All 12 outsider alterations caught."},{"name":"Python and TypeScript verifiers agree","value":"219 / 219","unit":null,"n":219,"split":"synthetic","note":null}],"dataset":"Synthetic notebooks from the vertical's test kit (entries, attachments by SHA-256, amendments, AI analyses with receipts, signatures, RFC 3161 tokens from an offline test TSA), attacked by an outsider and an insider holding the server key across 13 alterations in three configurations; plus one real recorded demo export with a FreeTSA token.","held_out":false,"caveats":["Synthetic notebooks only; no personal or real lab data.","No dev/test split: the verifier's rules were written first and not tuned to pass cases; only the attacker was changed after a run.","Against the operator, account signatures alone catch few changes: an insider can delete timestamp tokens that stop matching.","The quality of the AI analyses is not measured; the product records and receipts them, it does not grade them.","Real-world clock drift against the TSA and TSA certificate revocation are not measured."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/signed-lab-notebook"},"quality_evidence":[{"tier":"lite","label":"Lite · notebook, signatures and timestamps on CPU (self-host)","evidence":[{"metric":"Genuine exports verified","value":"300 / 300","source":"docs/evals/signed-lab-notebook.md"},{"metric":"Alterations caught, file edited without keys (13 kinds in 5 groups, 3 configurations)","value":"1,400 / 1,400","source":"docs/evals/signed-lab-notebook.md"},{"metric":"Alterations caught, insider with the server key, personal keys + TSA","value":"515 / 520 on the export alone; 520 / 520 with an earlier export","source":"docs/evals/signed-lab-notebook.md"},{"metric":"Verify a 10,000-entry notebook (17,107 chain entries)","value":"2.6 s Python; 4.7 s browser code (Node)","source":"docs/evals/signed-lab-notebook.md"}]},{"tier":"standard","label":"Standard · Qwen3.8-27B writes receipted AI analyses (hosted demo)","evidence":[{"metric":"AI analyses with a signed receipt covering the stored output","value":"5 / 5 end-to-end runs (smoke, recorder, self-host, two console analyses)","source":"scripts/smoke/signed-lab-notebook.py; docs/evals/signed-lab-notebook.md"},{"metric":"Insider edit of an AI output with a gateway receipt (real export)","value":"caught: the gateway receipt covers a different output","source":"docs/evals/signed-lab-notebook.md"},{"metric":"Sample analysis: outlier named, numbers checked","value":"named well F7 in the run recorded; 7 of 13 numbers in the data, 4 rounded, 2 flagged (one run, not an accuracy measure)","source":"docs/evals/signed-lab-notebook.md"},{"metric":"Tamper detection and 10k performance","value":"as the lite tier (same code)","source":"docs/evals/signed-lab-notebook.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass (pre-release server)","p50_ms":23700,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.00105},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory on our server, image built from docker/api/Dockerfile, api started with compose (named volume, python healthcheck), then the assemble prompt's smoke steps 1-9; torn down afterwards.","notes":"Ran against the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114 instead of the compose llm service; model-server startup not re-verified. Analysis receipt attested, FreeTSA token in 159 ms, export verified, altered export rejected. Host networking and port 8438 because other services held the default ports."},"known_limits":["Hosted check ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route): console flow with personal keys, own entry with a file, AI analysis, self-witness refused, timestamp rate limit, export, audit CSV, tamper buttons, verify page, Watch replay, Build example as written, 390 px layout. Production runs it once merged and deployed.","The server signs exports with its own key: an operator could rebuild a notebook that has only account signatures and no surviving TSA token. Personal keys and export copies kept by others catch that.","Signer identity is the name the key-holder sends unless the signer enrolled a personal key; no SSO or two-component sign-in (21 CFR 11.200) yet.","AI analyses are receipted, not graded; the number check only points at numbers not in the data.","FreeTSA is a free service with no SLA; TSA certificate revocation is not checked (the issuer is pinned).","POST /notebook/verify takes 8 MB by default; larger exports are checked in the browser (a 17 MB, 10,000-entry export took 4.7 s)."],"receipt_coverage":"full"},"cost_per_run_usd":0.00105,"rehearsal_bundle":{"url":"/samples/signed-lab-notebook.zip","checks":7,"bytes":6489},"models":[{"name":"decosa-api notebook module (decosa_api/verticals/notebook)","role":"Notebook: hash chain, members, amendments, e-signature rules, RFC 3161 client, export, audit CSV and verification (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Writes the AI analysis note from the selected entries and attached data files","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/signed-lab-notebook","page":"/tools/life-sciences/signed-lab-notebook","json":"/use-cases/signed-lab-notebook.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"green-claims-check","num":"45","name":"Green-claims substantiation check","status":"live","industries":["compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Verdict accuracy, v2 on test2 (held out, first run)","value":"43/44 (98%)","unit":null,"n":44,"split":"heldout","note":"v2.2 reruns on test2: 41/44 each."},{"name":"Planted violations flagged, v2 on test2 (held out)","value":"19/19","unit":null,"n":19,"split":"heldout","note":"v2.2 reruns: 18/19 each."},{"name":"False alarms on clean claims and non-claims, v2 on test2 (held out)","value":"0/25","unit":null,"n":25,"split":"heldout","note":"v2.2 reruns: 1/25 each."},{"name":"Verdict accuracy, v1 on test (held out)","value":"58/69 (84%)","unit":null,"n":69,"split":"heldout","note":"30/30 violations flagged, 5/39 false alarms; v2 was then changed after reading these errors."},{"name":"Verdict accuracy, word list only, test2","value":"12/44 (27%)","unit":null,"n":44,"split":"heldout","note":"Baseline without the model: 4/19 violations flagged."},{"name":"Model rewrites still using a generic or neutrality term, v2 on test2","value":"2/9","unit":null,"n":9,"split":"heldout","note":"Withheld from display since v2.1."}],"dataset":"18 synthetic marketing pieces with evidence files for fictional brands (dev 3, test 9, test2 6), each sentence labelled with the expected verdict and EU Empowering Consumers Directive rule.","held_out":true,"caveats":["Small and synthetic: 44 held-out sentences in 6 pieces, 2 to 4 planted per rule.","The same author (Claude, for Decosa) wrote the cases and the prompts, so they may share blind spots.","Prompts were changed after reading the v1 test errors, so v2 and later numbers on the test split are not held out; the shipped v2.3 on test2 is not held out either.","No regulator decision or court case was used as ground truth; labels are our reading of the Directive.","Text only: labels drawn as artwork, colours and imagery are not read.","The gateway is not fully deterministic at temperature 0; reruns differ by a sentence or two."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/green-claims-check"},"quality_evidence":[{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Held-out test2 (6 synthetic pieces, 44 sentences): verdicts right, first run / three reruns","value":"43/44 / 41/44 each","source":"docs/evals/green-claims-check.md (v2 and v2.2 result files), 25 Sep 2026"},{"metric":"Planted violations flagged, held out (test2 / v1 on test)","value":"19/19 (18/19 in reruns) / 30/30","source":"docs/evals/green-claims-check.md"},{"metric":"False alarms on clean claims and non-claims, held out (test2 / v1 on test)","value":"0/25 (1/25 in reruns) / 5/39","source":"docs/evals/green-claims-check.md"},{"metric":"Substantiated claims whose cited span holds the expected evidence (test2)","value":"16/16","source":"docs/evals/green-claims-check.md"},{"metric":"Per rule on test2 (2a, 4a, 4b, 4c, 10a, 6(2)(d)), first run","value":"precision and recall 1.00 each, on 3, 4, 2, 4, 3 and 2 planted claims; reruns: 10a recall 0.67, 2a precision 0.69","source":"docs/evals/green-claims-check.md"},{"metric":"Word list alone (no model), test2","value":"12/44 verdicts, 4/19 violations flagged","source":"docs/evals/green-claims-check/baseline-test2-eu.json"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":8600,"p95_ms":null,"runs":null,"receipts_per_run":26,"cost_per_run_usd":0.0126},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it, the compose file from this prompt, then its smoke steps against the already-running local Qwen3.8-27B vLLM on the direct route. All three samples ran (4.5-6 s each): Fernhollow banned_claims with 5 bans, Quillbrook nothing_flagged, Northwick (UK) high_risk; the report verified and a changed status failed. Model-server startup itself not re-verified."},"known_limits":["Text only: artwork, colours and label images are not read. Describe a label in words to have it checked.","Checks the Directive's text, not national transposing laws, and not the durability and repair bans (23d to 23j).","The eval is small and synthetic, and the gateway is not fully deterministic: the same piece can get a different verdict on a borderline claim from run to run. A person reviews every finding."],"receipt_coverage":"full"},"cost_per_run_usd":0.0126,"rehearsal_bundle":{"url":"/samples/green-claims-check.zip","checks":10,"bytes":3060},"models":[{"name":"decosa-api green module (decosa_api/verticals/green) on the promo pre-check engine, with the grounding module (decosa_api/verticals/grounding)","role":"Rulepack, sentences, evidence spans, verdicts, claim table and signed report (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Claim typing, grounding judge, evidence reader and rewrite","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/green-claims-check","page":"/tools/finance/green-claims-check","json":"/use-cases/green-claims-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"fi-disclosure-record","num":"46","name":"Auto F&I disclosure record","status":"live","industries":["sales-marketing","finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Per-check accuracy, test B (held out), run 1","value":"31/32","unit":null,"n":32,"split":"heldout","note":"Repeat run 32/32."},{"name":"Planted answers found, test B (held out), run 1","value":"8/8","unit":null,"n":8,"split":"heldout","note":null},{"name":"False alarms on clean calls, test B (held out), run 1","value":"1 of 14","unit":null,"n":14,"split":"heldout","note":"Repeat run 0 of 14."},{"name":"Per-check accuracy, test, run 2 (after two fixes)","value":"77/78","unit":null,"n":78,"split":"test","note":"Run 1, before the fixes: 76/78 with 2 of 30 false alarms."},{"name":"Per-check accuracy on ASR transcripts of TTS audio, run 1","value":"45/45","unit":null,"n":45,"split":"test","note":"Audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) and re-run: repeat 45/45; timestamps on the gold line 21/21 then 20/21 (first build: repeat 44/45, timestamps 21/21 then 19/21)."},{"name":"Timestamps on the gold line, test B, run 1","value":"15/15","unit":null,"n":15,"split":"heldout","note":"Repeat 15/16."}],"dataset":"17 synthetic F&I role-plays for a fictional dealer, written from SB 766 and Penal Code 632: dev 3, test 10 (4 clean, 6 planted), test B 4 written after test run 1 and never used to change anything; 5 also as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) through the diarizer.","held_out":true,"caveats":["Synthetic only: 17 short scripted conversations with clean TTS audio; no crosstalk, speed-read menus, other languages or 30 to 60 minute sessions.","The same author wrote the scripts, the expected answers and the prompts, so this is not an independent label set.","Test run 1 led to two fixes before the test set was re-run; test B is the held-out number.","Answer probabilities are not calibrated for F&I talk.","Cannot tell whether a written disclosure was clear and conspicuous, in the right language, or given in time."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/fi-disclosure-record"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out test B, 4 synthetic conversations: typed checks correct / planted problems found / mismatches found","value":"31/32 / 8/8 / 3/3","source":"decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25, gateway route, run 1; data and prompts written by the building agent, prompts frozen on a separate 3-conversation dev split"},{"metric":"test set, 10 conversations: typed checks correct (run 1 before two fixes / after)","value":"76/78 / 77/78; planted 12/13; mismatches 6/6","source":"decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25, gateway route"},{"metric":"false alarms on clean conversations (flag or review)","value":"test 0/30 after fixes (2/30 before); test B 1/14","source":"decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25"},{"metric":"5 conversations on real speech-recognition transcripts of synthetic audio: checks / planted / cited time within 3 s","value":"45/45 / 10/10 / 21/21 (repeat 45/45, 10/10, 20/21)","source":"decosa-api docs/evals/fi-disclosure-record.md (asr runs), measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: repeat 44/45, 9/10, 19/21), gateway route"},{"metric":"real F&I recordings reviewed by a dealer compliance officer","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":3300,"p95_ms":null,"runs":null,"receipts_per_run":10,"cost_per_run_usd":0.0027},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"Both smoke tests in the assembly prompt passed against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb4 flagged the cooling-off denial and the add-on never discussed, 10 attested receipts, record verified, 3.8 s; the pasted transcript flagged the service contract as presented as required at 00:05; no consent gave 400. Model-server startup and the diarizer path were not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route), driven from the branch site in headless Chromium, including at 390 px. The production API gets this vertical when the branch merges.","Measured on 17 synthetic role-plays written by the building agent, with clean TTS audio. It has not been measured on real F&I recordings (long sessions, crosstalk, Spanish) or with a dealer compliance officer's labels.","The Act's add-on and payment disclosures must be in writing. The tool reads them from the jacket as the dealer states them, and cannot judge whether they were clear, conspicuous or in the right language.","One systematic miss: a hint at cancelling ('if you did cancel there's a restocking fee') is read as explaining the 3-day right. The misstatement check still flags that call.","Audio intake (/fi/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.","It does not check advertising or the first written price (1784.41(a)), GAP loan-to-value (1784.42(a)(3)), or add-on payment timing (1784.42(b))."],"receipt_coverage":"full"},"cost_per_run_usd":0.0027,"rehearsal_bundle":{"url":"/samples/fi-disclosure-record.zip","checks":9,"bytes":2985},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Extraction (what was said about each add-on, prices, payment) and one typed yes/no/unclear check per question","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/fi-disclosure-record","page":"/tools/finance/fi-disclosure-record","json":"/use-cases/fi-disclosure-record.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"consented-dubbing","num":"48","name":"Consented creator dubbing","status":"live","industries":["entertainment","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Dub tracks exactly the video's length (sample count)","value":"11 of 11","unit":null,"n":11,"split":"synthetic","note":"Plus 6 of 6 synthetic frame-rate edge cases exact"},{"name":"Consent gate decisions as expected","value":"9 of 9","unit":null,"n":9,"split":"synthetic","note":"Deterministic code; all 9 receipted"},{"name":"ASR word error rate against the narration script","value":"3.9% (own-voice), 1.0% (dubber-voice)","unit":null,"n":null,"split":"synthetic","note":"Synthetic, clean narration; real creators will score worse"},{"name":"Glossary terms rendered as required","value":"98 of 98","unit":null,"n":98,"split":"synthetic","note":null},{"name":"Back-translation chrF against the source","value":"mean 73.4 (range 70.0-76.0)","unit":null,"n":null,"split":"synthetic","note":"A proxy for meaning kept, not a quality score"},{"name":"Lines cut short to fit","value":"GPU runs 12 of 160; CPU run 3 of 18","unit":null,"n":178,"split":"synthetic","note":null},{"name":"Planted subtitle errors caught","value":"793 of 793","unit":null,"n":793,"split":"synthetic","note":"16 kinds x 25 tries; shows each check fires, not real-world prevalence. 0 false alarms on the clean hand-made file"}],"dataset":"Two self-made narration videos (57 s and 43.2 s) of fictional creators with synthetic stock voices, 11 pipeline runs; a hand-written Spanish SDH file with planted errors; 6 synthetic length edge cases.","held_out":false,"caveats":["No human rating of the Spanish or of the voice: the numbers are measurable proxies.","Synthetic stock voices and clean narration; real voices depend on the consent clip's quality.","The subtitle checker and its planted errors have the same author.","The voice keeps some English accent; a native Spanish speaker has not rated it. No lip-sync, no stem separation, one speaker only."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/consented-dubbing"},"quality_evidence":[{"tier":"lite","label":"Lite · voice on CPU","evidence":[{"metric":"Voice real-time factor on CPU","value":"3.28","source":"measured on our server 2026-09-25, 1 run"},{"metric":"Exact video length","value":"1 of 1 run (3 of 18 lines cut short)","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"}]},{"tier":"standard","label":"Standard · the hosted demo, voice on a shared GPU","evidence":[{"metric":"Track length equals the video's to the sample","value":"11 of 11 runs; 6 of 6 synthetic edge cases (29.97 fps, audio longer or shorter than video, audio-only)","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Consent gate: decisions as expected","value":"9 of 9 (allowed, revoked, strike, territory, purpose, project, swapped voice, no entry), all receipted","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Subtitle QA: planted errors caught","value":"793 of 793 planted by script (16 kinds), 15 of 15 in the hand-made demo file, 0 false alarms on the clean hand-made file","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25; the planter and the checker have the same author"},{"metric":"ASR WER against the script","value":"1.0-3.9%","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Lines cut short to fit","value":"12 of 160 lines across 10 GPU runs","source":"decosa-api docs/evals/consented-dubbing.md, 2026-09-25"},{"metric":"Dub quality rated by a native speaker","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":74000,"p95_ms":null,"runs":null,"receipts_per_run":43,"cost_per_run_usd":0.0026},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image plus the documented voice-runtime layer, compose with named volumes, pointed at the already-running local Voxtral and Qwen3.8-27B (direct route), voice on CPU; then torn down.","notes":"The revoked sample was refused; the own-voice dub reached review in 337 s (including the first download of the voice weights) with an exact 2,736,000-sample track, glossary 10 of 10 and a render receipt; approval gave a C2PA-signed release (state Valid) and a record that verifies; no transcript or approver text in the container logs. Found on the way: chatterbox-tts pins numpy<1.26, which has no wheels for the image's Python 3.12, so the documented layer now builds the voice venv with Python 3.11 via uv."},"known_limits":["One speaker, English to Spanish; no lip-sync; music under the voice is not carried into the dub track.","The voice keeps some English accent (cross-language cloning), and no native speaker has rated it yet: that is why approval is required.","About 1 line in 13 is cut short to fit its slot (12 of 160 on GPU runs).","Approval proves the job's key or session pressed Approve on the exact draft, not which person did.","C2PA credentials use a development certificate, so public validators show the issuer as untrusted."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0026,"rehearsal_bundle":{"url":"/samples/consented-dubbing.zip","checks":17,"bytes":774333},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Transcribes each speech span of the source, and listens back to every dubbed line","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewer","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Chatterbox Multilingual (t3_23lang)","role":"Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermark","license":"MIT","hf_repo":"ResembleAI/chatterbox"},{"name":"ECAPA-TDNN speaker embeddings (ONNX export)","role":"Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (tool 47)","license":"Apache-2.0","hf_repo":"speechbrain/spkrec-ecapa-voxceleb"}],"licence":"permissive","links":{"metrics":"/metrics/consented-dubbing","page":"/apps/consented-dubbing","json":"/use-cases/consented-dubbing.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"disclosure-preflight","num":"49","name":"Synthetic-performer disclosure and S&P pre-flight","status":"live","industries":["entertainment","sales-marketing"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted script issues found (full config)","value":"27 of 29 (93%)","unit":null,"n":29,"split":"test","note":"Dev: 31 of 32"},{"name":"Precision of script flags (full config)","value":"1.00 (27 flags, 0 false)","unit":null,"n":27,"split":"test","note":null},{"name":"Clean scripts with a flag","value":"0 of 8","unit":null,"n":8,"split":"test","note":"Without the yes/no step: 2 of 8; word lists only: 2 of 8"},{"name":"Word lists only (no model): planted issues found","value":"7 of 29","unit":null,"n":29,"split":"test","note":null},{"name":"OCR reads the AI label on finished Decosa ads / false alarm on unlabelled shots","value":"18 of 18 frames / 0 of 60 frames","unit":null,"n":78,"split":"synthetic","note":"Own label style on own footage only"}],"dataset":"48 synthetic ad scripts for fictional brands (24 dev, 24 test), with issues planted from fixed pools that share nothing between dev and test; Decosa's own UGC ad renders and ten raw H3 shots for the label check.","held_out":true,"caveats":["The plants are blatant, one clear instance each, in short synthetic scripts; expect lower recall and precision on real scripts.","The same author wrote the eval script, the planted pools and the prompts; one prompt change was made after the first dev run.","The label check measures Decosa's own label style on its own footage; third-party labels may not be read.","Conspicuousness under NY GBL 396-b is a legal judgement and is not measured. No synthetic-performer detector is used or claimed."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/disclosure-preflight"},"quality_evidence":[{"tier":"lite","label":"Lite · code only, any CPU","evidence":[{"metric":"Planted script issues found / precision / clean scripts flagged (test, 29 plants in 24 synthetic scripts, 8 clean)","value":"7 of 29 / 0.78 / 2 of 8","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (word lists only: profanity 4 of 4, music markup 3 of 6; the false flags are generic music cues)"},{"metric":"AI label read on Decosa ads / read back after burn-in / false label on unlabelled clips","value":"18 of 18 / 60 of 60 / 0 of 60 frames","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (3 finished ads, 10 raw MiniMax H3 shots)"},{"metric":"C2PA marking valid with content intact and the disclosure assertion","value":"10 of 10","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25"}]},{"tier":"standard","label":"Standard · adds script checks by Qwen3.8-27B (hosted demo)","evidence":[{"metric":"Planted script issues found / precision / clean scripts flagged (test)","value":"27 of 29 (93%) / 1.00 (27 flags, 0 false) / 0 of 8","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (dev run twice, one change to the rating question between runs; test run once)"},{"metric":"By category (test): real person / brand / profanity / rating trigger / music","value":"4 of 4 / 7 of 8 / 4 of 4 / 7 of 7 / 5 of 6","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (misses: 'Apple' and the song 'Happy')"},{"metric":"Without the yes/no step (test): found / precision / clean scripts flagged","value":"28 of 29 / 0.93 / 2 of 8","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25"},{"metric":"Demo samples giving the expected performers, rules, consent states and flags","value":"7 of 7","source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25"}]}],"benchmark":{"title":"How well does it catch planted issues, and does the label stick?","intro":"48 short synthetic ad scripts for fictional brands (written by Qwen3.8-27B, receipted, read by hand) had real names, other brands, swear words, rating triggers and hit songs planted in them, from pools split between dev and test; 8 scripts per split carried only hard negatives. The label and marking were checked on Decosa's own renders.","rows":[{"label":"Planted script issues found (test, 29)","value":"93% (27)","detail":"27 flags, none false; 0 of 8 clean scripts flagged"},{"label":"Without the typed yes/no step","value":"28 found, 2 false","detail":"both false flags were generic music cues"},{"label":"Word lists only (no model)","value":"7 of 29","detail":"profanity and music markup only"},{"label":"Label read back after burn-in","value":"60 of 60 frames","detail":"and 0 of 60 frames of unlabelled clips read as labelled"},{"label":"C2PA marking valid, content intact","value":"10 of 10","detail":null}],"points":[{"heading":"Where it fails","text":"Single words that are also everyday words: 'Apple' in 'I cancelled Apple for this', and the song 'Happy'. The yes/no step that removes generic music cues also dropped that one real song."},{"heading":"What this does not show","text":"The plants are blatant, one instance each, in short synthetic scripts. The label check read our own label style on our own footage. No real-world recall, and no judgement of what counts as 'conspicuous', is claimed."}],"source":"decosa-api docs/evals/disclosure-preflight.md, 2026-09-25"},"verification":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":3600,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0012},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone into a clean directory, docker build (50 s), the api service with named volumes, a development C2PA certificate from scripts/provenance_devcert.py, pointed at the running local vLLM (Qwen3.8-27B) over host networking; then torn down.","notes":"Verified on 2026-09-25: the image has ffmpeg, Tesseract 5.5.0, the font and c2pa-python; the planted script gives the same five flags and consent states as the hosted run in 2.7 s (five attested calls); the raw clip gets the label (6 of 6 frames read back) and a valid C2PA marking in 5.3 s; the attestation verifies and fails when one decision is changed; no script text or names in the logs. The Decosa-ad samples are hosted renders and show as unavailable, as expected. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Advisory, not a legal clearance; 'conspicuous' is not measured.","Synthetic performers are found only from provenance or a declaration.","Script recall is measured on short synthetic scripts with blatant plants.","Logos and faces in the picture are not checked yet.","Consent comes from the consent ledger (tool 47) by id; the hosted demo's id_demo identities are fictional. A server without the ledger reports real performers as unchecked, never cleared."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0012,"rehearsal_bundle":{"url":"/samples/disclosure-preflight.zip","checks":16,"bytes":260023},"models":[{"name":"decosa-api disclosure module (decosa_api/verticals/disclosure)","role":"The pre-flight: provenance lookup, performer evidence, rules for NY and the EU, consent lookups by ledger id, the profanity and music-cue word lists, the report, sign-off and the signed attestation (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Tesseract OCR 5 (English)","role":"Reads six frames for an AI label (whole frame, then the top and bottom bands), before and after the label is added","license":"Apache-2.0","hf_repo":null},{"name":"c2pa-python 0.37 (c2pa-rs)","role":"Reads the file's C2PA credential and signs the marking: the source as a parentOf ingredient, c2pa.opened and c2pa.edited actions with the IPTC digital source type, and an ai.decosa.disclosure assertion","license":"MIT OR Apache-2.0","hf_repo":null},{"name":"TrustMark Q (decoder)","role":"Decodes Decosa's invisible video watermark (TrustMark), so a Decosa render whose credential was stripped is still recognised","license":"MIT","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: lists candidate names, brands, profanity, rating triggers and music cues in the script, then answers a typed yes/no for each (typed-judgment, calibrated probability)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/disclosure-preflight","page":"/tools/media/disclosure-preflight","json":"/use-cases/disclosure-preflight.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"animatic-studio","num":"50","name":"Script to animatic","status":"live","industries":["entertainment","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Shot-list coverage of lines needing a shot, before any code repair","value":"100% in 10 of 10 runs","unit":null,"n":10,"split":"synthetic","note":"5 scripts x 2 runs, 78 required lines per pass"},{"name":"Invented dialogue","value":"0","unit":null,"n":10,"split":"synthetic","note":"0 by construction: dialogue is cited by line id and inserted by code"},{"name":"Grounding verdicts on shot actions: supported / partial / unsupported / contradicted","value":"128 / 17 / 3 / 3","unit":null,"n":151,"split":"synthetic","note":"Of the 6 hard flags read by hand, 2 were real and 4 over-strict"},{"name":"Planted invented actions judged unsupported","value":"10 of 10","unit":null,"n":10,"split":"synthetic","note":"66 unplanted shots in the same pass: 57 supported, 9 partial, 0 unsupported or contradicted"},{"name":"Character identity (CLIP cosine to sheet), unlocked vs locked look (default)","value":"0.528 vs 0.690","unit":null,"n":18,"split":"synthetic","note":"Nearest-sheet accuracy 0.33 vs 0.75 (12 frames; chance about 0.42)"},{"name":"Model cost per shot list","value":"$0.0035-0.013","unit":null,"n":null,"split":"synthetic","note":"At the gateway list price; render GPU time not billed in the demo"}],"dataset":"5 scripts: 3 original CC0 (screenplay, 30 s ad, game cutscene) and 2 public-domain stage plays (Wilde, Glaspell); 18 single-character shots for the consistency test.","held_out":false,"caveats":["No held-out split: the Last Lamp sample was used during development, and prompt rules were added after a first run on it.","The planted inventions are blunt on purpose, and the plants and the checker have the same author; a subtle invention lands in partial at best.","Whole-frame CLIP scores are confounded by background and framing: read them as differences between arms, not absolute truth.","Frames are not checked against existing characters; one robot came out close to a well-known film robot and was withdrawn. Review for resemblance before sharing.","A shot-list prompt rule and a judge fix went in after these numbers and are not re-measured."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/animatic-studio"},"quality_evidence":[{"tier":"lite","label":"Lite · breakdown only, no GPU renderer","evidence":[{"metric":"Model coverage of script lines","value":"100% in 10 of 10 runs","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"},{"metric":"Planted invented events flagged by the grounding check","value":"10 of 10","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · the hosted demo","evidence":[{"metric":"Script lines covered by the model's shot list, before code repair","value":"100% in 10 of 10 runs (5 scripts, 78 required lines)","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"},{"metric":"Invented dialogue in the cut","value":"0 by construction (the model cites line ids; code inserts the words); 0 quoted phrases outside the cited lines in 10 runs","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"},{"metric":"Planted invented events flagged by the grounding check","value":"10 of 10 (unsupported)","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"},{"metric":"Character consistency (CLIP ViT-L/14 similarity of each frame to its character sheet)","value":"0.69 with each look repeated in every frame vs 0.53 without; nearest-sheet identification 0.75 vs 0.33 (18 frames, 3 scripts). Reference-image conditioning scored 0.81 but copied the sheet's pose, so it is off.","source":"decosa-api docs/evals/animatic-studio.md, 2026-09-26"},{"metric":"Render time per full animatic on the shared GPU","value":"359-726 s for 7-18 shots (6 runs); about 13 s per frame when the GPU is free; 45-47 s to re-render one shot","source":"measured on our server 2026-09-26"}]},{"tier":"best","label":"Best · self-host with LTX-2.3 or MiniMax H3 (licence pending) motion","evidence":[{"metric":"Any","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":531000,"p95_ms":null,"runs":null,"receipts_per_run":16,"cost_per_run_usd":0.0069},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image plus the Kokoro layer from the assemble prompt, compose with named volumes, pointed at the already-running local Qwen3.8-27B (direct route) and ComfyUI; the rehearsal bundle; then torn down.","notes":"9 of 9 rehearsal checks passed (render 576 s, C2PA stamped, record verified), no script text in the container logs. Found on the way: frames could not be moved from the work volume to the data volume (cross-device rename); fixed with a regression test."},"known_limits":["Frames are rough and characters are only partly consistent from shot to shot (no identity adapter).","Frames aren't checked against existing characters; review for resemblance before sharing. A vague look can drift towards a familiar character (in testing, a one-line robot description came out close to a well-known film robot).","Long speeches stretch their shots past the drafted length; the timeline shows by how much.","The grounding check also marks harmless staging details as partial; a person decides.","Renders share one GPU with other demos: several minutes per animatic, longer when the queue is busy.","C2PA credentials use a development certificate, so public validators show the issuer as untrusted."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0069,"rehearsal_bundle":{"url":"/samples/animatic-studio.zip","checks":9,"bytes":1704},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Drafts the shot list (line ids, framing, action, duration) and a look per character; the same model is the grounding judge that checks each shot against its lines","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Wan2.2-VACE-Fun-A14B","role":"Character sheets, one still per shot (text only: every frame repeats each visible character's look), and optional image-to-video motion from a shot's still","license":"Apache-2.0","hf_repo":"alibaba-pai/Wan2.2-VACE-Fun-A14B"},{"name":"Wan2.2-Lightning T2V 4-step LoRAs","role":"4-step distillation LoRAs for Wan2.2 (used at 6 steps, cfg 2 for stills)","license":"Apache-2.0","hf_repo":"lightx2v/Wan2.2-Lightning"},{"name":"Kokoro-82M","role":"Temp dialogue in stock voicepacks, only for speakers the consent ledger allows","license":"Apache-2.0","hf_repo":"hexgrad/Kokoro-82M"},{"name":"c2pa-rs via c2pa-python","role":"C2PA content credential on the MP4, with the consent-ledger links of the voiced speakers","license":"MIT OR Apache-2.0","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/animatic-studio","page":"/apps/animatic-studio","json":"/use-cases/animatic-studio.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"audio-drama-studio","num":"51","name":"Audio drama and narrated story studio","status":"live","industries":["entertainment","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Radio scripts: speaker right (code alone)","value":"100% (389/389)","unit":null,"n":389,"split":"test","note":"Spoken lines found: also 100% (389/389)"},{"name":"Sound cue mapped to the right library tag (model)","value":"99.0% (96/97)","unit":null,"n":97,"split":"test","note":"Keyword code alone: 92.8% (90/97)"},{"name":"Cue with nothing in the library left as \"no sound\"","value":"87.5% (7/8)","unit":null,"n":8,"split":"test","note":"Varies run to run: a second pass missed three such cues"},{"name":"Prose: speaker right","value":"96.8% (209/216)","unit":null,"n":216,"split":"test","note":"Dev: 94.2% (65/69)"},{"name":"Public-domain excerpts: speaker right","value":"100% (30/30)","unit":null,"n":30,"split":"test","note":"Adapted excerpts of The Red-Headed League and The Monkey's Paw"},{"name":"Word error rate of the finished episodes (machine transcription)","value":"0.6%-5.0%","unit":null,"n":4,"split":"synthetic","note":"4 sample episodes, 175-625 words each"},{"name":"Parse cost","value":"about $0.0007 per parse","unit":null,"n":62,"split":"test","note":"62 receipted calls on the test split, $0.044 at the gateway list price"}],"dataset":"Seeded synthetic radio scripts (8 dev, 30 test) and prose stories (8 dev, 30 test), plus hand-labelled public-domain excerpts (Poe as dev; Doyle and Jacobs adaptations as test); 4 rendered sample episodes for loudness and listening checks.","held_out":true,"caveats":["Mostly synthetic data from a generator written by the same author as the prompts; three prompt changes were made after looking at dev.","The two test excerpts were written from memory, so they are adaptations, not exact texts.","The eval's author cannot hear: acting, how effects land and music fit were not checked by ear. A human listen is needed before publishing.","Results vary run to run on the shared gateway even at temperature 0 (the \"no sound\" row)."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/audio-drama-studio"},"quality_evidence":[{"tier":"lite","label":"Lite · CPU only, radio scripts","evidence":[{"metric":"Speakers right on held-out radio scripts (code alone)","value":"100% (389/389)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Cue mapped to the right library sound by keywords","value":"92.8% (90/97)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · the hosted demo","evidence":[{"metric":"Prose speaker attribution (held-out)","value":"96.8% planted (209/216); 100% public domain (30/30)","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Cue mapping (held-out)","value":"99.0% (96/97); music cues 26/26","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"},{"metric":"Loudness spec met on delivered files","value":"all sample episodes and chapter files","source":"measured on our server 2026-09-26"},{"metric":"Word error rate heard back by ASR","value":"0.6-5.0% on 4 sample episodes","source":"decosa-api docs/evals/audio-drama-studio.md, 2026-09-26"}]},{"tier":"best","label":"Best · compose new music per episode","evidence":[{"metric":"Music prompt guard on held-out prompts","value":"precision 100%, recall 98%","source":"decosa-api docs/evals/music-gen-cleared.md, 2026-09-25"},{"metric":"Composed cues in this studio","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":28000,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.001},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image and the docker/drama voice layer, compose with named volumes and host networking, pointed at the running local Qwen3.8-27B (direct route); rehearsal bundle; then torn down.","notes":"Builds took 24 s and 100 s (warm cache). The rehearsal passed 16 of 16 checks in 20.7 s, including an 18 s render on 8 CPU threads: parse, an out-of-project performer refused, the house cast allowed, loudness in spec, C2PA with consent links, and the signed record verified. No speaker model in the image, so the voice check said not run."},"known_limits":["Kokoro voices are clear but flat: deliveries change pace and level only. British voices are Kokoro's weakest.","Prose attribution is about 97% right on held-out stories: check the parse before rendering.","ACX does not accept AI narration; the audiobook mode meets the technical spec only.","C2PA credentials use a development certificate, so public validators show the issuer as untrusted.","Composing new music needs the studio GPU queue; the hosted demo uses the cue library."],"receipt_coverage":"partial"},"cost_per_run_usd":0.001,"rehearsal_bundle":{"url":"/samples/audio-drama-studio.zip","checks":16,"bytes":3134},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Kokoro-82M","role":"Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPU","license":"Apache-2.0","hf_repo":"hexgrad/Kokoro-82M"},{"name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","role":"Score: theme and sting cues rendered through the music-gen-cleared path (tool 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPU","license":"MIT","hf_repo":"ACE-Step/Ace-Step1.5"},{"name":"decosa-api drama module (decosa_api/verticals/drama) + FFmpeg","role":"Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"ECAPA-TDNN speaker embeddings (ONNX export)","role":"Consent ledger's speaker check (tool 47): does each role's rendered voice match the voice enrolled in its entry?","license":"Apache-2.0","hf_repo":"speechbrain/spkrec-ecapa-voxceleb"}],"licence":"permissive","links":{"metrics":"/metrics/audio-drama-studio","page":"/apps/audio-drama-studio","json":"/use-cases/audio-drama-studio.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"music-video-studio","num":"52","name":"Music video from your track","status":"live","industries":["music","entertainment"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Lyric lines within 0.3 s / 1 s of human timing","value":"81.8% / 87.7%","unit":null,"n":6,"split":"test","note":"6 held-out songs; 4 of 6 songs have at least 80% of lines within 1 s. Baseline 16.8% within 1 s. Dev: 67.0% / 67.9%."},{"name":"Beat F-measure (±70 ms), artist's BPM given","value":"0.868","unit":null,"n":40,"split":"test","note":"Without BPM: 0.591. Drums-only grooves: an upper bound for full mixes."},{"name":"Typed check says no to a swapped (mismatched) scene","value":"67.3%","unit":null,"n":52,"split":"test","note":"The code checks (citations, anchor words) carry the grounding; the typed check lets a third through."},{"name":"Anchor words found in the cited lines","value":"96.2%","unit":null,"n":52,"split":"test","note":"Citations valid: 100%"},{"name":"Rendered cut offset from the beat grid (median / max)","value":"7.0-10.0 ms / 16.3 ms","unit":null,"n":null,"split":"synthetic","note":"5 renders of one excerpt; against the detected grid, not a human one. Some cuts between similar shots are not detectable (13-21 of 20-21 found)."},{"name":"Render cost, one format","value":"about $0.32 per minute of video","unit":null,"n":null,"split":"synthetic","note":"At an assumed $1.69/h GPU rental price; the hosted demo does not bill renders."}],"dataset":"JamendoLyrics MultiLang (9 English songs with human word/line timings: 3 dev, 6 test), Groove MIDI Dataset (40 test grooves; tuned on the validation split), and renders of one CC BY excerpt and one generated track.","held_out":true,"caveats":["Visual quality is not scored; 480p upscaled clips are clearly softer than closed models.","Beat tracking measured on synthesised drums only, not real full mixes.","English only (the aligner is English-only).","Small sets: 6 test songs; the cut-detector threshold was set on the same render it was measured on.","The containerised analyzer placed one beat grid about 160 ms later than on the host."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/music-video-studio"},"quality_evidence":[{"tier":"lite","label":"Lite · timed captions and a cited treatment, no GPU for video","evidence":[{"metric":"Lyric line starts within 0.3 s / 1 s of human timing (6 held-out CC BY-ND songs)","value":"81.8% / 87.7% (baseline 16.8% within 1 s)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Beat F-measure (±70 ms), 40 human-played grooves","value":"0.87 with the artist's BPM, 0.59 without","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; drums only"},{"metric":"Treatment citations valid / anchor words found (52 held-out scenes)","value":"100% / 96%","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · the hosted demo, clips on one 96 GB card","evidence":[{"metric":"Cuts found in the rendered files / offset from the beat grid","value":"133 of 145 planned cuts found, 0 false; median 7-10 ms, max 16.3 ms (half a frame at 30 fps)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; 7 files from 5 renders; offsets against the detected beat grid"},{"metric":"GPU time per minute of video","value":"about 680 GPU-seconds for one format, 1,360-1,740 for both","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26; shared card"},{"metric":"Disclosure rules on the delivered files (tool 49)","value":"label read on 6 of 6 sampled frames and the C2PA marking valid in all 4 files checked","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Typed grounding check: swapped (mismatched) scenes caught","value":"67% (the code checks on citations and anchor words do the rest)","source":"decosa-api docs/evals/music-video-studio.md, 2026-09-26"},{"metric":"Visual quality","value":"not measured; 480p upscaled, clearly below closed models","source":null}]},{"tier":"best","label":"Best · MiniMax H3 clips (licence pending), self-host","evidence":[{"metric":"clip quality","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","evidence":[{"metric":"render time per clip against one card with offload","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":15100,"p95_ms":null,"runs":null,"receipts_per_run":7,"cost_per_run_usd":0.0021},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image and services/mvideo, compose with named volumes on host networking, pointed at the already-running local Qwen (direct route) and ComfyUI; then torn down.","notes":"No rights statement: 403; the sample analysed in 19 s (112.35 BPM, 18 lines, 6 grounded scenes, 20 cuts, 7 attested receipts); a 9:16 render finished with a C2PA credential, the disclosure rules met and a record that verified. Found on the way: the containerised analyzer's beat grid sat about 160 ms later than the host's on the same file (different decoder), and a clip of flickering neon fooled the cut detector (fixed: isolated spikes only)."},"known_limits":["Visuals are 480p clips upscaled to 720p: a stylised visualiser, clearly below closed video models and the H3 samples on this site.","Lyric timing is English only; 82% of held-out lines start within 0.3 s, and when it slips whole passages slip.","Without the artist's BPM the beat tracker got the tempo right on 43% of test grooves (100% with it); which beat is \"one\" is a heuristic.","The hosted demo caps a video at 60 s and renders share one GPU: 3 renders per session, 3 per API key per day, about 11 minutes per minute of video per format.","The rights statement is not verified; the clearance pre-check covers only a small open catalogue.","C2PA credentials are signed by a development CA: valid signature, untrusted issuer in public validators."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0021,"rehearsal_bundle":{"url":"/samples/music-video-studio.zip","checks":10,"bytes":844310},"models":[{"name":"decosa-mvideo-analyze (services/mvideo)","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","license":"AGPL-3.0-or-later (decosa-api)","hf_repo":null},{"name":"wav2vec2-large-960h-lv60-self","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","license":"Apache-2.0","hf_repo":"facebook/wav2vec2-large-960h-lv60-self"},{"name":"Qwen3.8-27B (NVFP4)","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa-api clearance module (decosa_api/verticals/clearance)","role":"Sample and lyric clearance pre-check on the upload (tool 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs","role":"Clips: text-to-video, one 5 s clip per scene per format","license":"Apache-2.0","hf_repo":"alibaba-pai/Wan2.2-VACE-Fun-A14B"},{"name":"decosa-api mvideo module + FFmpeg + c2pa-python","role":"The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (tool 49) on the file; C2PA credential per file; the signed record","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/music-video-studio","page":"/apps/music-video-studio","json":"/use-cases/music-video-studio.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"sar-narrative-desk","num":"53","name":"SAR narrative desk","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted wrong numbers caught in reference narratives, first run (no model)","value":"1,598 of 1,600 (99.9%)","unit":null,"n":1600,"split":"test","note":"1,600 of 1,600 after a fix made on seeing the two misses; 0 of 1,920 correct numbers flagged."},{"name":"Wrong numbers planted in the model's own sentences caught (v3)","value":"340 of 356 (95.5%)","unit":null,"n":356,"split":"test","note":"Amounts 133/133, dates 165/174, counts 34/40."},{"name":"Invented sentences held in investigator drafts","value":"18 of 20","unit":null,"n":20,"split":"test","note":"The other two were flagged (partial), so all 20 were caught. Re-run 28 Sep 2026 on the same weights (direct route) after the judge was given the institution-conclusions and place-name context; the same day before that change: 19 held, 1 flagged; 26 Sep (gateway): 19 held, 1 flagged."},{"name":"True reference sentences held / flagged","value":"0 of 172 held; 19 of 172 (11%) flagged","unit":null,"n":172,"split":"test","note":"28 Sep 2026 re-run (direct route, same weights). Before the change, same day: 0 held, 27 flagged. 26 Sep (gateway): 1 held, 32 flagged."},{"name":"Numbers the model wrote that the checker flagged (v3 drafts)","value":"0 of 533","unit":null,"n":533,"split":"test","note":"4 of 299 sentences held, one of them a false hold by the judge."},{"name":"Invented facts in randomly sampled checked sentences (manual read, v2)","value":"0 of 70; 2 of 70 (3%) minor unsupported details passed","unit":null,"n":70,"split":"dev","note":"Read by the building agent only; v2 is partly development data."}],"dataset":"Synthetic case files from the vertical's own generator (fictional people, accounts and bank). Seeds 1-20 used for building, 1000+ for the eval: 200 cases for the numeric check, three draft sets of 15 runs (v1, v2, v3; code changed after v1 and v2, v3 run after the last checker change), 20 reference narratives with one invented sentence each.","held_out":true,"caveats":["Everything is synthetic, written by the same agent that wrote the checker and the prompts: evidence the mechanisms work on this generator, not accuracy on real bank data.","Hallucinations were judged by the building agent reading the sentences; no second reader.","Code was changed after v1 and v2, so those sets are partly development data; two small fixes after v3 are not re-measured.","Comparator claims and percentages are not planted in the numeric check, and reference narratives are template prose.","The screen's thresholds were set on this generator; 5 of 40 clean cases were flagged for structuring."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/sar-narrative-desk"},"quality_evidence":[{"tier":"lite","label":"Lite · numbers and citations only, no GPU","evidence":[{"metric":"Correct numbers flagged, reference narratives (200 held-out synthetic cases)","value":"0 of 1,920","source":"docs/evals/sar-narrative-desk.md, part A, 26 Sep 2026"},{"metric":"Planted wrong numbers caught (amounts, dates, counts), reference narratives","value":"1,600 of 1,600 (1,598 before a bug fix)","source":"docs/evals/sar-narrative-desk.md, part A"},{"metric":"Planted wrong numbers caught in the model's own sentences","value":"340 of 356 (95.5%): amounts 133/133, dates 165/174, counts 34/40","source":"docs/evals/sar-narrative-desk.md, part D (v3)"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Drafted sentences traced / flagged / held (v3: 15 synthetic cases, 299 sentences)","value":"272 / 23 / 4 (1 of the 4 a false hold)","source":"docs/evals/sar-narrative-desk.md, part B"},{"metric":"Numbers the model wrote that failed the check (v3)","value":"0 of 533","source":"docs/evals/sar-narrative-desk.md, part B"},{"metric":"Citation validity (v3)","value":"0 uncited sentences; 1 unknown id of 931","source":"docs/evals/sar-narrative-desk.md, part B"},{"metric":"Invented facts in investigator drafts held / flagged (20 drafts)","value":"18 / 2 (all 20 caught); 0 of 172 true sentences held, 19 flagged (28 Sep re-run)","source":"docs/evals/sar-narrative-desk.md, part C"},{"metric":"Typology coverage: planted typology found by the screen and named in the draft (12 cases, v3)","value":"12/12 and 12/12; none named that the screen did not find","source":"docs/evals/sar-narrative-desk.md, part B"},{"metric":"Hallucinated facts in traced sentences (manual read, 70 sampled)","value":"0 invented facts; 2 small unsupported details (a state, 'international')","source":"docs/evals/sar-narrative-desk.md, part E"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":24700,"p95_ms":null,"runs":null,"receipts_per_run":30,"cost_per_run_usd":0.017},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it with DECOSA_SAR_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The planted-errors check held exactly the three planted sentences (1.7 s); the structuring draft ran in 7.9 s with 51 of 51 numbers traced; the report verified, a changed status failed, and receipts were attested. Model-server startup itself not re-verified."},"known_limits":["Synthetic only on the hosted demo. Every number here comes from our own synthetic generator; it has not been run on real case files.","A date is checked for membership in the cited rows, and tied to its amount when the sentence pairs them; a date moved onto another cited day with no amount beside it can pass. Counts can match another subset of the cited rows.","The screen's thresholds are ours. It flagged structuring on 5 of 40 clean cash-business cases.","It checks that what is written is traced to the case file, not that nothing is missing, and it sees nothing outside the case file (watch lists, other institutions, prior SARs).","The ledger must be CSV; PDF statements are not read."],"receipt_coverage":"full"},"cost_per_run_usd":0.017,"rehearsal_bundle":{"url":"/samples/sar-narrative-desk.zip","checks":15,"bytes":4197},"models":[{"name":"decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)","role":"Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Section drafting and the grounding judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/sar-narrative-desk","page":"/tools/finance/sar-narrative-desk","json":"/use-cases/sar-narrative-desk.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"claims-conduct-pack","num":"54","name":"Insurance claims-file conduct pack","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted problems found, all checks","value":"52 of 54 (precision 0.98, recall 0.96)","unit":null,"n":54,"split":"test","note":"48 test files, run once after the prompts were frozen"},{"name":"False flags","value":"1","unit":null,"n":48,"split":"test","note":"A shortened flood exclusion called a misrepresentation"},{"name":"Clean files with any flag","value":"0 of 12","unit":null,"n":12,"split":"test","note":"Dev 0 of 7, replies 0 of 6, samples 0 of 2"},{"name":"Held-out phrasing files: found / false / missed","value":"26 / 0 / 0","unit":null,"n":24,"split":"heldout","note":"Sentences never seen while the prompts were written; denial-reason wordings are not held out"},{"name":"Timeliness items with the right start, act and deadline","value":"122 of 122","unit":null,"n":122,"split":"test","note":"Dev: 64 of 65"},{"name":"Late replies found (targeted replies set)","value":"6 of 6, 0 false flags","unit":null,"n":12,"split":"test","note":null}],"dataset":"Synthetic claim files for a fictional insurer from scripts/claims_cases.py: dev 24 files, test 48 (half with held-out phrasing), a targeted replies set of 12, and 6 demo samples; problems planted on a structured truth.","held_out":true,"caveats":["Synthetic, templated files: real claim files are longer and messier (scanned letters, email chains, several claimants). These numbers do not predict accuracy on a carrier's files.","The same author wrote the generator, the planted problems and the prompts.","Small positive counts for some checks: decision (2), review notice (2) and lowball (3) on the test set.","Denial-reason wordings are the same six in every set, so misquote detection is not held out on wording.","Rules are a subset (total-loss valuation, subrogation and others are not checked); no OCR in this build."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/claims-conduct-pack"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"planted problems flagged / false alarms","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out test, 48 synthetic files run once: planted problems flagged / false flags","value":"52/54 / 1","source":"decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route; prompts frozen on a separate 24-file dev set; half the test files use phrasing never seen while writing the prompts (26/26 found there)"},{"metric":"clean files with any flag","value":"0/12 test, 0/7 dev, 0/6 replies, 0/2 samples","source":"decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route"},{"metric":"date accuracy on the test set: events with the right date / timeliness items with the right start, act and deadline","value":"247/247 / 122/122","source":"decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route"},{"metric":"per check on the test set (found/planted)","value":"acknowledgement 13/13, decision 2/2, payment 4/4, denial reason 11/12, review notice 2/2, misrepresentation 10/11 (1 false), low offer 3/3, investigation 7/7; late replies 6/6 on a 12-file targeted set","source":"decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route"},{"metric":"AI use recorded, and whether a person was involved read right","value":"31/31","source":"decosa-api docs/evals/claims-conduct-pack.md, measured on our server 2026-09-26, gateway route"},{"metric":"real, de-identified claim files reviewed by a claims-quality auditor","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"planted problems flagged / false alarms","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"planted problems flagged / false alarms","value":"not measured yet","source":null}]}],"benchmark":{"title":"How well does it do on synthetic claim files?","intro":"84 synthetic files from a generator with planted problems (late acknowledgement, a clause misquoted, no review notice, a low offer, no investigation and more), across the NAIC, California and Texas rulepacks. Prompts were written on a 24-file dev set; the 48-file test set was run once.","rows":[{"label":"Planted problems flagged, test set","value":"52 of 54","detail":"1 false flag; the miss came back as review"},{"label":"Clean files with any flag","value":"0 of 27","detail":"test, dev, replies and demo sets together"},{"label":"Deadlines right end to end","value":"122 of 122","detail":"start date, act date and deadline, test set"},{"label":"Cost per file","value":"about $0.004","detail":"5 model calls, 9,351 tokens on the demo file, gateway list price"}],"points":[{"heading":"Where it fails","text":"Both errors are the model reading a letter against the policy: a letter that widened 'livery conveyance' to 'any business purpose' came back as review, not flag, and a letter that shortened the flood exclusion was called a misrepresentation."},{"heading":"What it does not show","text":"The files are templated and synthetic. Real files are longer and messier (scanned letters, email chains, several claimants). Measure it on your own closed files before relying on it."}],"source":"decosa-api docs/evals/claims-conduct-pack.md, 26 Sep 2026"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":10100,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.004},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): the planted file flagged ack, reason:1 (quoting 'organized race or speed contest'), review_notice and misrepresent, 5 attested receipts, record verified, 5.0 s; the clean file had 0 flags; sign-off pointed at the first record; the rehearsal bundle passed 8/8. Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured on 84 synthetic, templated files written by the building agent; not on real claim files or with a claims auditor's labels.","Three rulepacks only (NAIC model, California, Texas). The NAIC pack is a baseline, not any state's law.","Business days skip weekends and US federal holidays; deadlines on weekends are not rolled forward, and one or two days late on such a deadline is marked review.","The reviewer's name on a sign-off is as given; identity is not verified."],"receipt_coverage":"full"},"cost_per_run_usd":0.004,"rehearsal_bundle":{"url":"/samples/claims-conduct-pack.zip","checks":8,"bytes":3830},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the file (dated events, denial reasons, amounts, AI mentions, each with a quote), grounds each denial reason in the policy, and answers the typed conduct questions","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/claims-conduct-pack","page":"/tools/finance/claims-conduct-pack","json":"/use-cases/claims-conduct-pack.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"collections-call-qa","num":"55","name":"Collections and servicing call QA","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Per-rule accuracy, test run 1","value":"81/83","unit":null,"n":83,"split":"test","note":"Run 2: 81/83"},{"name":"Per-rule accuracy, test B run 1 (written after test run 1)","value":"32/33","unit":null,"n":33,"split":"heldout","note":null},{"name":"Planted answers found, test run 1","value":"11/12","unit":null,"n":12,"split":"test","note":"Test B: 3/3; audio: 8/9"},{"name":"False alarms on clean calls: test / test B run 1","value":"0 of 48 / 2 of 21","unit":null,"n":69,"split":"test","note":"Test B run 2: 1 of 21"},{"name":"Call-log findings (hand-labelled, first runs)","value":"61/61","unit":null,"n":61,"split":"test","note":"Test, test B and audio"},{"name":"Per-rule accuracy on ASR transcripts of TTS audio, run 1","value":"48/49","unit":null,"n":49,"split":"test","note":"Audio re-voiced with Decosa house voices (Kokoro-82M) and re-run: accuracy unchanged from the earlier macOS-voice audio. Timestamps on the gold line 19/21 (mean error 0.53 s, max 4.2 s); run 2: 20/21"}],"dataset":"17 synthetic scripted calls written from 12 CFR part 1006 and 1024.39-41 with fictional companies and people: 3 dev, 10 test, 4 test B, plus 6 of them as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) run through the diarizer.","held_out":true,"caveats":["The same author wrote the scripts, the labels, the prompts and the log code; the labels are one reading of the rule text.","17 short scripted calls with clean TTS audio; real calls (accents, crosstalk, Spanish, long calls), voicemails and limited-content messages are not measured.","The call log is taken as given; presumptions depend on facts outside it.","Answer probabilities come from a confidence table fit on another domain and are not calibrated for collection calls."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/collections-call-qa"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out test B, 4 synthetic calls: typed checks correct / planted problems found / call-log findings","value":"32/33 / 3/3 / 11/11","source":"decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route, run 1; data, labels and prompts written by the building agent, prompts frozen on a separate 3-call dev split"},{"metric":"test set, 10 calls: typed checks correct / planted found / call-log findings","value":"81/83 / 11/12 / 33/33 (repeat run the same, plus one unlabelled log flag)","source":"decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route"},{"metric":"false alarms on clean calls (flag or review)","value":"test 0/48; test B 2/21 (repeat 1/21)","source":"decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route"},{"metric":"6 calls on real speech-recognition transcripts of synthetic audio: checks / planted / log findings / cited time within 3 s","value":"48/49 / 8/9 / 17/17 / 20/21","source":"decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route (asr runs)"},{"metric":"real collection calls reviewed by a compliance QA lead","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"typed checks correct / planted problems found","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":4150,"p95_ms":null,"runs":null,"receipts_per_run":9,"cost_per_run_usd":0.0026},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the sheriff threat at 00:27, 7-in-7 (c8) and calling hours (c1), 9 attested receipts, record verified, 3.4 s; the pasted 21:30 New York call flagged the threat and the hours; no consent gave 400. One prompt bug found and fixed: the log-only step fed back the normalised call log, which the API does not accept as input. Model-server startup and the diarizer path were not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured on 17 synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real collection calls (accents, Spanish, long calls, voicemails) or with an independent reviewer's labels.","One systematic miss: an agent who names themselves a debt collector before confirming who answered is not flagged as revealing the debt before identity (t6 on every run).","The call-log findings are the rule's presumptions computed from the log given. Letters, texts, emails, consent given elsewhere and an attorney's response are outside it unless added as events. Days are counted in the consumer's first time zone.","Audio intake (/collections/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled."],"receipt_coverage":"full"},"cost_per_run_usd":0.0026,"rehearsal_bundle":{"url":"/samples/collections-call-qa.zip","checks":12,"bytes":3310},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Extraction (what the called person asked for or said, with the line) and one typed yes/no/unclear check per QA question","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/collections-call-qa","page":"/tools/finance/collections-call-qa","json":"/use-cases/collections-call-qa.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"incident-notification-pack","num":"56","name":"Incident notification pack","status":"live","industries":["compliance-trust","software"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted problems caught, held-out scenarios","value":"20 / 20","unit":null,"n":20,"split":"test","note":"Two scenarios run twice: wrong times, wrong counts, unsupported and contradicted claims, removed elements, stale figures, a late 8-K."},{"name":"Planted problems caught, dev scenarios","value":"28 / 28","unit":null,"n":28,"split":"dev","note":null},{"name":"Deadlines right, hand-labelled cases, first run","value":"35 / 44","unit":null,"n":44,"split":"test","note":"Labelled from the rules by a separate agent; the 9 misses were one bug, fixed, after which 44 / 44 (not independent)."},{"name":"Clean sentences held, held-out scenarios","value":"1 / 42","unit":null,"n":42,"split":"test","note":"3 of 42 flagged as partly supported."},{"name":"False contradictions on clean held-out packs","value":"2 / 4","unit":null,"n":4,"split":"test","note":"One pattern: a later time read as the detection time."},{"name":"Model-drafted sentences held (real errors caught)","value":"2 / 108","unit":null,"n":108,"split":"synthetic","note":"Both were times copied from the wrong entry; 9 flagged; 28 / 28 required elements given."}],"dataset":"Five synthetic incidents written by the building agent (ransomware at a SaaS vendor, a public bucket at a DORA payment institution, an exploited router vulnerability under the CRA; held out: a DDoS on a DNS provider and email compromise at a billing company), each run clean and with planted problems, twice; plus 44 deadline cases hand-labelled by a separate agent.","held_out":true,"caveats":["Everything is synthetic, written by the same agent that wrote the checker and the prompts; small n (5 scenarios).","Plants are single clear errors; legal adequacy and subtle understatement are not measured.","The deadline set was used to find and fix a bug, so 44 / 44 after the fix is not held out.","The judge varies run to run: the same clean sentence was traced in one run and flagged or held in another.","Hand labels are by an AI agent from the rules text, not by counsel."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/incident-notification-pack"},"quality_evidence":[{"tier":"lite","label":"Lite · clocks and timeline, no GPU","evidence":[{"metric":"Deadlines right, 44 hand-labelled cases (SEC, NIS2, DORA, CRA, California)","value":"44 / 44 after one bug fix; 35 / 44 first run","source":"docs/evals/incident-notification-pack.md, part 1, 26 Sep 2026"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Planted problems caught in supplied notices (2 repeats)","value":"28 / 28 dev; 20 / 20 held-out test","source":"docs/evals/incident-notification-pack.md, parts 2-6"},{"metric":"Unsupported or contradicted claims held","value":"10 / 10","source":"docs/evals/incident-notification-pack.md, parts 2-6"},{"metric":"Clean sentences held / flagged","value":"1 / 102 held; 14 / 102 flagged","source":"docs/evals/incident-notification-pack.md, part 7"},{"metric":"False contradictions on clean packs","value":"0 of 6 dev packs; 2 of 4 test packs (one pattern)","source":"docs/evals/incident-notification-pack.md, part 7"},{"metric":"Model-drafted sentences traced / flagged / held (5 scenarios)","value":"97 / 9 / 2 (both held were real errors)","source":"docs/evals/incident-notification-pack.md, part 8"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":23320,"p95_ms":null,"runs":null,"receipts_per_run":16,"cost_per_run_usd":0.006},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged to main) into a clean directory, the api image built from docker/api/Dockerfile, compose api service with a named data volume, DECOSA_INCIDENT_SYNTHETIC_ONLY=0, direct route to the already-running local Qwen3.8-27B vLLM (network_mode host instead of starting a second model server). The ransomware sample without the synthetic flag: both planted sentences held, impact missing, 3 contradictions, the NIS2 early warning late, report and timeline record verified, a changed status failed, receipts attested, 3.8 s; the CRA sample with a drafted final report 4.4 s. Torn down after. Model-server startup itself not re-verified."},"known_limits":["Synthetic only on the hosted demo, and every number here comes from our own synthetic incidents.","Coverage says a traced sentence addresses an element, not that it says enough.","The contradiction check can read a later time as the detection time (seen in 2 of 4 clean held-out packs).","Member State bank holidays, national NIS2 formats and other US states are not modelled.","The clocks are only as right as the tags: the team marks when it became aware, classified or determined materiality."],"receipt_coverage":"full"},"cost_per_run_usd":0.006,"rehearsal_bundle":{"url":"/samples/incident-notification-pack.zip","checks":13,"bytes":4802},"models":[{"name":"decosa-api incident pack (decosa_api/verticals/incident) with the numeric block (decosa_api/verticals/numeric), the dates module (decosa_api/verticals/claims/dates.py) and the grounding module (decosa_api/verticals/grounding)","role":"Timeline, hash chain, deadline clocks, citation, time and number checks, element coverage, cross-notice consistency, signed pack and sign-off (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Notice drafting, the grounding judge, and the review call (element coverage and quoted facts)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/incident-notification-pack","page":"/tools/finance/incident-notification-pack","json":"/use-cases/incident-notification-pack.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"filing-tieout","num":"57","name":"Filing tie-out and MD&A grounding","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"False flags on untouched held-out 10-K MD&As (frozen rules)","value":"11 in 6,039 figures (1.8 per 1,000)","unit":null,"n":6039,"split":"test","note":"13 flags in 6 of 40 filings; 2 were real inconsistencies (stale note numbers in Harmonic). 6 flags, 2 real, after post-test fixes (not a clean test)."},{"name":"Note table that disagrees with a statement, caught","value":"79/79","unit":null,"n":79,"split":"test","note":null},{"name":"Scale error (million/billion) caught","value":"88/142 (62%)","unit":null,"n":142,"split":"test","note":null},{"name":"Another period's figure caught","value":"38/150 (25%)","unit":null,"n":150,"split":"test","note":null},{"name":"Flipped direction caught in code","value":"38/130 (29%)","unit":null,"n":130,"split":"test","note":null},{"name":"Wrong percentage change caught","value":"14/91 (15%)","unit":null,"n":91,"split":"test","note":null},{"name":"Changed figure caught","value":"22/153 (14%)","unit":null,"n":153,"split":"test","note":"Figures that tied in the clean pass; over all MD&A amounts 10/156."},{"name":"Flipped direction claims shown as a claim mismatch","value":"20/40","unit":null,"n":40,"split":"test","note":"3 of the same 40 sentences unflipped were also shown as mismatches; flips the judge called contradicted but was not sure of go to review. Measured 30 Sep on the direct route (same weights)."},{"name":"MD&A figures tied or computed","value":"34.6%","unit":null,"n":6039,"split":"test","note":null}],"dataset":"FY2025 10-Ks of US large accelerated filers from the SEC Financial Statement Data Sets 2026q1, picked by hash of the accession number and fetched from EDGAR: 100 for development (two rounds) and 40 fetched after the rules were frozen (test). Errors planted in code in the real MD&A text; the untouched MD&A is the clean set.","held_out":true,"caveats":["One agent wrote the rules, planted the errors and judged which flags on clean filings were real errors.","The planted errors are ours; real draft errors may differ.","Recall is low by design; most figures are not in the tagged tables at all and come back untraced.","The claim-judge eval forced the judge on every sentence; the product's sentence selection changed afterwards.","MD&A tables and 10-Q quarters are not measured."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/filing-tieout"},"quality_evidence":[{"tier":"lite","label":"Lite · figures only, no GPU","evidence":[{"metric":"False flags on untouched held-out 10-K MD&As (40 filings, frozen rules)","value":"11 in 6,039 figures (1.8 per 1,000); 2 more flags were real errors","source":"docs/evals/filing-tieout.md, 26 Sep 2026"},{"metric":"Planted errors caught, held-out: note table vs statement / scale / period / direction / % change / changed figure","value":"79/79 / 88/142 / 38/150 / 38/130 / 14/91 / 22/153","source":"docs/evals/filing-tieout.md, 26 Sep 2026"},{"metric":"MD&A figures tied or computed (held-out)","value":"34.6%","source":"docs/evals/filing-tieout.md, 26 Sep 2026"},{"metric":"Parser vs SEC's own extracted statement figures","value":"10,808 of 10,813","source":"docs/evals/filing-tieout.md, 26 Sep 2026"}]},{"tier":"standard","label":"Standard · one GPU for claims in words (hosted demo)","evidence":[{"metric":"Flipped direction claims shown as mismatches (40 held-out sentences)","value":"20/40","source":"docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026"},{"metric":"True sentences shown as mismatches (the same 40, unflipped)","value":"3/40","source":"docs/evals/filing-tieout/claims-results-p1-1.json, 30 Sep 2026"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":9453,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.0004},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_TIEOUT_PUBLIC_ONLY=0, on the direct route to the running local Qwen3.8-27B. Re-run on commit 2562f1c with the rehearsal bundle (16 of 16 checks). The planted sample caught all seven plants in 0.21 s (0.61 s with one claim read), the clean sample had no flags, the record verified and a changed status failed, and an EDGAR fetch worked from the container."},"known_limits":["Only MD&A prose is tied; tables inside MD&A are not.","Recall is low by design: most wrong figures in single-figure sentences come back untraced rather than flagged.","Figures for segments, non-GAAP measures or narrower scopes that share a line's label can still be flagged against the consolidated line.","Drafts must be iXBRL or pasted text with CSV tables; Word and PDF are not read.","10-Q quarter periods are handled but were not evaluated."],"receipt_coverage":"full"},"cost_per_run_usd":0.0004,"rehearsal_bundle":{"url":"/samples/filing-tieout.zip","checks":16,"bytes":3482},"models":[{"name":"decosa-api tie-out (decosa_api/verticals/tieout) with the numeric grounding block's table-cell matcher (decosa_api/verticals/numeric/cells.py)","role":"iXBRL parser, figure tie-out, period and scale rules, direction and cross-reference checks, workpaper and signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Claims in words (optional): the grounding judge reads a sentence against the named lines' figures","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/filing-tieout","page":"/tools/finance/filing-tieout","json":"/use-cases/filing-tieout.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"sanctions-disposition-record","num":"58","name":"Sanctions alert disposition record","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Same-party pairs proposed as false positive (the risky direction)","value":"0 of 1,094","unit":null,"n":1094,"split":"test","note":"Synthetic customers built from real list entries; dev 0 of 1,102."},{"name":"False-positive clearance precision","value":"635 / 635 (100%)","unit":null,"n":635,"split":"test","note":"Share of false-positive proposals that were different parties."},{"name":"Same-party pairs proposed as true match","value":"863 (78.9%)","unit":null,"n":1094,"split":"test","note":"The other 231 were held as needs more information (name only, day/month swapped, renewed passport)."},{"name":"Per-field status accuracy (labelled fields)","value":"99.77% of 2,199","unit":null,"n":2199,"split":"test","note":null},{"name":"Model reading agrees with the code proposal","value":"71 of 81 (88%)","unit":null,"n":81,"split":"test","note":"6 reviews that first lost both calls to a gateway outage were re-run after the fix; the outage is not counted."},{"name":"Rationale sentences passing the code check","value":"278 of 278 (100%)","unit":null,"n":278,"split":"test","note":"Manual read of 30: 0 invented facts."},{"name":"Planted rationale errors held by the checker","value":"2,776 of 2,776 (100%)","unit":null,"n":2776,"split":"test","note":"531 of 531 correct sentences passed after one fix made on seeing this set (516 before)."}],"dataset":"Synthetic customers against real OFAC, EU and UK list entries (snapshot of 26 Sep 2026), labelled by construction: 13 same-party and 11 different-party cases, split by list-entry uid into dev (1,990 pairs) and test (1,967 pairs); 81 test reviews through the hosted model (6 re-run after a gateway outage).","held_out":true,"caveats":["The customers, the variants and the rules were all written by the same agent: this shows the rules do what they say on list-derived data, not accuracy on a real alert queue.","A true match with wrong customer data (a mistyped date of birth or ID) is not in the set and would be proposed as a false positive.","The checker's planted errors are templates; the model made no error the checker caught, so its recall on real model errors is unknown.","Names in scripts other than Latin and Cyrillic are not compared.","The manual read had one reader, the building agent."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/sanctions-disposition-record"},"quality_evidence":[{"tier":"lite","label":"Lite · the comparison and the proposal, no GPU","evidence":[{"metric":"Same-party pairs proposed as false positive (held-out test split)","value":"0 of 1,094","source":"docs/evals/sanctions-disposition-record.md, part A, 26 Sep 2026"},{"metric":"False-positive clearance precision (test)","value":"635 of 635 (100%)","source":"docs/evals/sanctions-disposition-record.md, part A"},{"metric":"Field status accuracy on labelled fields (test)","value":"99.77% of 2,199","source":"docs/evals/sanctions-disposition-record.md, part A"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Typed reading agrees with the code proposal (81 test reviews)","value":"71 of 81 (88%); it never proposed clearing a same-party pair","source":"docs/evals/sanctions-disposition-record.md, part B"},{"metric":"Rationale sentences passing the code check","value":"278 of 278","source":"docs/evals/sanctions-disposition-record.md, part B"},{"metric":"Planted rationale errors held by the checker (wrong year, country, ID, polarity, citation)","value":"2,776 of 2,776; 531 of 531 correct sentences passed","source":"docs/evals/sanctions-disposition-record.md, part C"},{"metric":"Invented facts in a manual read of 30 sampled rationale sentences","value":"0","source":"docs/evals/sanctions-disposition-record.md, part B"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":8380,"p95_ms":null,"runs":null,"receipts_per_run":2,"cost_per_run_usd":0.0006},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from docker/api/Dockerfile with DECOSA_SANCTIONS_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The dob-mismatch review took 0.86 s (false_positive by R-DISQ, both receipts attested, 3 grounded sentences); the record verified and an edited one failed; an override without a reason and a risky clear without a second reviewer both got 422. The documented list refresh ran in the container in 41 s. Model-server startup itself not re-verified."},"known_limits":["Synthetic customers only on the hosted demo, and every number here comes from our own generator on real list entries; not yet run on a real alert queue.","A true match whose customer data is wrong (a mistyped date of birth or ID) can be proposed as a false positive; the analyst checks the source document.","Names in scripts other than Latin and Cyrillic are not compared, so those alerts are held.","It reviews one alert against one list entry. It does not screen, and it does not cover ownership (the 50 Percent Rule), licences or list changes after the snapshot.","Records are not stored: you keep them for 10 years."],"receipt_coverage":"full"},"cost_per_run_usd":0.0006,"rehearsal_bundle":{"url":"/samples/sanctions-disposition-record.zip","checks":13,"bytes":2133},"models":[{"name":"decosa-api sanctions desk (decosa_api/verticals/sanctions), with the typed-judgment core (vertical 24) and the session hash chain (record, vertical 07)","role":"List snapshot, field comparison, proposal rules, rationale check, signed record, audit sample (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Independent typed reading and the rationale","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/sanctions-disposition-record","page":"/tools/finance/sanctions-disposition-record","json":"/use-cases/sanctions-disposition-record.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"medicare-call-record","num":"59","name":"Medicare sales-call record","status":"live","industries":["healthcare","sales-marketing"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Per-rule accuracy, test B run 1 (never used to change anything)","value":"40/40","unit":null,"n":40,"split":"heldout","note":"Run 2: 40/40"},{"name":"Per-rule accuracy, test run 4","value":"101/101","unit":null,"n":101,"split":"test","note":"Run 1 (before the one code change): 101/101"},{"name":"Planted typed answers found, test run 4","value":"9/9","unit":null,"n":9,"split":"test","note":"Test B: 2/2; audio: 9/9 (also on the re-voiced audio)"},{"name":"Call-sheet items, test run 4 / test B run 1","value":"79/79 / 31/32","unit":null,"n":111,"split":"test","note":"Test run 1, before the needs change: 77/79"},{"name":"Benefit claims graded right (good ok, bad flagged), test run 4","value":"25/25 / 4/4","unit":null,"n":29,"split":"test","note":"Test B: 8/8 / 2/2; audio (re-voiced 26 Sep 2026 with Decosa house voices) runs 1 and 2: 18/19 / 2/3, both misses on t9 where the diarizer split a line and the scorer paired claims with the wrong gold lines (first build: 19/19 / 3/3 and 18/19 / 3/3)"},{"name":"False alarms on clean calls, test run 4 / test B run 1","value":"3 of 87 / 0 of 42","unit":null,"n":129,"split":"test","note":"All 3 are claims the plan facts do not mention"},{"name":"Cited time inside the spoken gold line, audio run 1","value":"28/29","unit":null,"n":29,"split":"test","note":"Audio re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 21/29 within 3 s of the line's start; run 2: 29/29. First build (macOS voices): 29/29, 16/29 within 3 s"}],"dataset":"17 synthetic scripted Medicare sales calls with fictional agency, plans and people, written from 42 CFR 422/423 subpart V: 3 dev, 10 test, 4 test B, plus 8 of them as TTS audio (Decosa house voices, Kokoro-82M, each allowed by the consent ledger) run through the diarizer.","held_out":true,"caveats":["The same author wrote the scripts, the labels, the prompts and the code; the labels are one reading of the rule text.","17 short scripted calls with clean TTS audio; real sales calls (long, accents, transfers, Spanish) are not measured.","One code change (the premiums needs topic) was made after test run 1, so test is not fully held out for that item; test B was never used to change anything. The scorer's audio matching was refined after the audio runs, before this write-up.","Benefit claims are judged against the plan facts given, not the plan's filed benefits.","Answer probabilities and claim confidences come from tables fit on other domains and are not calibrated for sales calls."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/medicare-call-record"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"typed checks correct / claims graded right","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out test B, 4 synthetic calls: typed checks / call-sheet items / claims graded right / false alarms on the 2 clean calls","value":"40/40 / 31/32 / 10/10 / 0 of 42","source":"decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 1 (run 2 the same); prompts frozen on a 3-call dev split; data, labels and prompts written by the building agent"},{"metric":"test set, 10 calls: typed checks / planted typed answers / call-sheet items / good claims ok / bad claims flagged","value":"101/101 / 9/9 / 79/79 / 25/25 / 4/4","source":"decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4, after one code change made on test run 1 (77/79 call-sheet items before it)"},{"metric":"false alarms on the 4 clean test calls (flag or review)","value":"3 of 87 (claims the plan facts do not mention)","source":"decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26, gateway route, run 4"},{"metric":"on diarized TTS audio, 8 calls: typed checks / call-sheet items / claims graded right / cited time inside the spoken line","value":"81/81 / 63/63 / 20/22 / 28/29 (run 2: 20/22 claims, 29/29)","source":"decosa-api docs/evals/medicare-call-record.md, measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 22/22 and 21/22 claims, 29/29), gateway route, asr runs 1 and 2"}]},{"tier":"best","label":"Best · two 96 GB cards","evidence":[{"metric":"typed checks correct / claims graded right","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"typed checks correct / claims graded right","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":7468,"p95_ms":null,"runs":null,"receipts_per_run":15,"cost_per_run_usd":0.0058},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the disclaimer timing, \"free\", the final-expense pitch and the $3,000 dental claim (contradicted), 15 attested receipts, record verified, retention 2029-11-05 / 2032-11-05, 3.8 s; the pasted cold call flagged Medicare, free, unsolicited, no SOA and the dental claim; no consent gave 400."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured on 17 short synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real sales calls (20-60 minutes, accents, transfers, Spanish) or with an independent reviewer's labels.","Claims the plan facts do not mention come back flagged even when true (2-3 on a clean drug-plan call); give the full Summary of Benefits.","ASR errors become claim errors: \"eyewear\" heard as \"in-store\" was flagged in one audio run.","Audio intake (/medicare/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled."],"receipt_coverage":"full"},"cost_per_run_usd":0.0058,"rehearsal_bundle":{"url":"/samples/medicare-call-record.zip","checks":12,"bytes":4075},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Two extractions (products, first benefit discussion, enrollment steps; the agent's benefit claims), one typed yes/no/unclear check per question, and one grounding judgment per benefit claim against the plan facts","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"}],"licence":"permissive","links":{"metrics":"/metrics/medicare-call-record","page":"/clinics/medicare-call-record","json":"/use-cases/medicare-call-record.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"denial-appeal-packet","num":"60","name":"Claim denial appeal packet","status":"live","industries":["healthcare","finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Recommendation right, fresh held-out set","value":"36 of 40","unit":null,"n":40,"split":"test","note":"test2, run once after the date check moved into code; 4 misses said don't appeal where gold says get documentation first"},{"name":"Unsupported cases with no appeal and no letter","value":"26 of 26","unit":null,"n":26,"split":"test","note":"test2; the earlier test set: 21 of 22"},{"name":"Supported cases with an appeal and a letter","value":"11 of 11","unit":null,"n":11,"split":"test","note":"test2; the earlier test set: 14 of 15"},{"name":"Recommendation right, first test set","value":"36 of 40","unit":null,"n":40,"split":"test","note":"Run once before the date check; one wrong appeal from a date misread"},{"name":"Criteria status right per policy requirement","value":"126 of 133","unit":null,"n":133,"split":"test","note":"test2; test 125 of 133; every requirement was found in the policy"},{"name":"Deadline right (next level and date)","value":"80 of 80","unit":null,"n":80,"split":"test","note":"test and test2; gold computed by separate code"},{"name":"Notice date read from the denial","value":"40 of 40","unit":null,"n":40,"split":"test","note":"cases where the request did not give it"},{"name":"Kept letter sentences with an unsupported clinical fact","value":"0 of 245","unit":null,"n":245,"split":"test","note":"Read by the building agent; 2 kept sentences overstated the policy"},{"name":"Held-out phrasing cases, recommendation right","value":"34 of 40","unit":null,"n":40,"split":"heldout","note":"Sentences never seen while the prompts were written: test 18 of 20, test2 16 of 20 (template phrasing: 18 of 20 and 20 of 20)"}],"dataset":"Synthetic denials, chart excerpts and policies from scripts/appeal_cases.py: dev 18 cases, test 40 and test2 40 (half with held-out phrasing), plus 5 demo samples; CMS NCD excerpts (public domain) and an invented commercial policy.","held_out":true,"caveats":["Synthetic, templated cases written by the same author as the prompts; real charts and commercial policies are longer and messier. These numbers do not predict accuracy on a provider's denials.","The unsupported-sentence count is the building agent's own reading of the letters, not an independent rater's.","After the first test run, date windows were moved into code; the fresh test2 set measures that change. Later small changes were checked on the demo samples only.","Four policies and five denial types; the gold for undocumented versus not met is a judgement call in a few chart phrasings.","Federal deadline rules only; state programs, plan documents and provider contracts are not modelled."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/denial-appeal-packet"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"recommendation and criteria accuracy","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"fresh held-out set (test2, 40 synthetic cases, run once): recommendation right","value":"36/40; all 4 misses said don't appeal where the gold says get documentation first","source":"decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26, gateway route; prompts frozen on an 18-case dev set; half the cases use phrasing never seen while writing the prompts"},{"metric":"planted unsupported cases (test2 / test): no appeal and no letter","value":"26/26 / 21/22 (the one miss: the model read a 53-day gap as 23; date windows are now checked in code, which test2 measures)","source":"decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26"},{"metric":"supported cases (test2 / test): appeal with a letter","value":"11/11 / 14/15","source":"decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26"},{"metric":"criteria status right, per policy requirement (test2 / test)","value":"126/133 / 125/133; every requirement was found in the policy (133/133)","source":"decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26"},{"metric":"deadline (next level and date) right / notice date read from the denial","value":"80/80 / 40/40","source":"decosa-api docs/evals/denial-appeal-packet.md, measured on our server 2026-09-26; gold deadlines computed by separate code with hard-coded holidays"},{"metric":"letter sentences kept after the check that state a clinical fact the chart doesn't support (read by the building agent)","value":"0/245; 2 kept sentences overstated the policy (said PSG is required where the NCD lists alternatives); 32 of 260 draft sentences were cut or marked [CHECK] by the check","source":"decosa-api docs/evals/denial-appeal-packet.md, test and test2 letters, 2026-09-26"},{"metric":"real, de-identified denials rated by a denials specialist","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"recommendation and criteria accuracy","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"recommendation and criteria accuracy","value":"not measured yet","source":null}]}],"benchmark":{"title":"Does it say no when it should?","intro":"98 synthetic cases from a generator with a structured truth: Medicare CPAP, cochlear-implant and bariatric denials against CMS coverage determinations, lumbar MRI denials against an invented payer policy, and billing denials. Some charts support every criterion; others contradict one or leave one out. Prompts were written on an 18-case dev set; a 40-case test set was run once, a date check was added in code, and a fresh 40-case set was run once.","rows":[{"label":"Unsupported cases with no appeal and no letter","value":"47 of 48","detail":"test2 26/26, test 21/22"},{"label":"Supported cases with an appeal and a letter","value":"25 of 26","detail":"test2 11/11, test 14/15"},{"label":"Deadlines right","value":"80 of 80","detail":"next level and date, test and test2"},{"label":"Cost per packet with a letter","value":"about $0.014","detail":"21 model calls, 34,908 tokens on the CPAP demo case, gateway list price"}],"points":[{"heading":"Where it fails","text":"It reads a criterion the chart leaves out as not met rather than undocumented (6 of 8 wrong recommendations across both sets): the advice becomes don't appeal where it should be get the documentation first. Once, before the date check moved into code, it read a 53-day gap as 23 and recommended an appeal the chart did not support."},{"heading":"What it does not show","text":"The cases are templated and synthetic, written by the same author as the prompts. Real charts are longer, scanned and messier, and commercial policies are denser than these excerpts. Measure it on your own closed denials before relying on it."}],"source":"decosa-api docs/evals/denial-appeal-packet.md, 26 Sep 2026"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":12900,"p95_ms":null,"runs":null,"receipts_per_run":21,"cost_per_run_usd":0.0139},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): deadline 2027-01-11, cpap-appeal appeal with a letter in 11.9 s (21 attested receipts, AHI criterion quoting 11.2), cpap-no-appeal don't appeal with no letter, admin-missing-npi fix the claim, record verified; the rehearsal bundle passed 10/10. Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured on 98 synthetic, templated cases written by the building agent; not on real denials or with a denials specialist's labels.","Federal deadline rules only; state external review, plan documents and provider contracts can differ. Deadlines other than external review are not moved off weekends.","Criteria come only from the policy text pasted in; non-covered indications are not mapped, and nested alternatives (an option with its own list of options) are read as one option.","Changes after the test2 run (a template opening line, stray reference markers stripped, alternatives worded as accepted options, source titles in the number guard) were checked on the demo samples only."],"receipt_coverage":"full"},"cost_per_run_usd":0.0139,"rehearsal_bundle":{"url":"/samples/denial-appeal-packet.zip","checks":10,"bytes":5680},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the denial (notice date, reasons, service), splits the payer policy into criteria, checks each criterion against the chart, drafts the letter when the chart supports it, and judges every letter sentence (the grounding judge)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/denial-appeal-packet","page":"/clinics/denial-appeal-packet","json":"/use-cases/denial-appeal-packet.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"kids-content-preflight","num":"61","name":"Made-for-kids content pre-flight","status":"live","industries":["creative-media","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Made for kids or not (full check)","value":"39 / 40","unit":null,"n":40,"split":"test","note":"Dev: 24 / 24. The miss was a close call (p = 0.59)."},{"name":"Audience, 4-way (primary / mixed / general / mature)","value":"37 / 40","unit":null,"n":40,"split":"test","note":"Two mixed-audience uploads called primary"},{"name":"Setting mismatches found","value":"15 / 15, 0 false","unit":null,"n":15,"split":"test","note":null},{"name":"Made for kids or not, audience-only audit","value":"39 / 40","unit":null,"n":40,"split":"test","note":"5 calls per upload instead of 15"},{"name":"Planted requests for children's details flagged","value":"3 / 3","unit":null,"n":3,"split":"test","note":null},{"name":"Frame check: picture-only stills, vision on","value":"10 / 12","unit":null,"n":12,"split":"heldout","note":"6 / 12 with vision off. Public-domain images labelled by the building agent; never used for tuning"}],"dataset":"64 synthetic uploads (24 dev, 40 test) written from the FTC's factors, YouTube's made-for-kids guidance and the Disney complaint, with the uploader's setting and planted data requests and sales pitches; 12 public-domain stills (6 children's-book illustrations, 6 not for children) for the frame check.","held_out":true,"caveats":["The same author wrote the uploads, the labels and the prompts; one prompt change was made on dev before the test split ran.","The test split ran twice: the first run overlapped a gateway outage (one upload unanswered; otherwise the same scores) and was re-run unchanged; the complete run is reported.","Short synthetic uploads that are on the nose; real channels, long vlogs and non-English videos will do worse.","The frame check used historical illustrations, not modern kids' video. Not a COPPA compliance determination."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/kids-content-preflight"},"quality_evidence":[{"tier":"lite","label":"Lite · audience only (the catalogue audit)","evidence":[{"metric":"held-out test, 40 synthetic uploads: made for kids or not / setting mismatches found","value":"39 / 40 / 15 of 15, 0 false","source":"decosa-api docs/evals/kids-content-preflight.md, measured on our server 2026-09-26, gateway route"},{"metric":"planted requests for children's details / sales pitches flagged (word lists only)","value":"2 of 3 / 0 of 3","source":"decosa-api docs/evals/kids-content-preflight.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · full check on text (hosted demo)","evidence":[{"metric":"held-out test, 40 synthetic uploads: made for kids or not / 4-way audience","value":"39 / 40 / 37 / 40","source":"decosa-api docs/evals/kids-content-preflight.md, measured on our server 2026-09-26, gateway route; prompts frozen on a separate 24-upload dev split"},{"metric":"setting mismatches found / false mismatches / close calls on correct uploads","value":"15 of 15 / 0 / 2","source":"decosa-api docs/evals/kids-content-preflight.md, 2026-09-26"},{"metric":"planted requests for children's details / sales pitches flagged","value":"3 of 3 / 2 of 3","source":"decosa-api docs/evals/kids-content-preflight.md, 2026-09-26"},{"metric":"demo samples: audience and setting check as planted","value":"7 of 7","source":"decosa-api docs/evals/kids-content-preflight.md, 2026-09-26"}]},{"tier":"best","label":"Best · adds the frame check (vision)","evidence":[{"metric":"12 public-domain stills, picture only: made-for-kids call agrees with the label, vision off / on","value":"6 of 12 / 10 of 12","source":"decosa-api docs/evals/kids-content-preflight.md, measured on our server 2026-09-26, gateway route (children's-book illustrations 0 of 6 / 4 of 6; other stills 6 of 6 both ways)"},{"metric":"frame check on real kids' video","value":"not measured yet","source":null}]}],"benchmark":{"title":"How often is the made-for-kids call right, and are mismatches caught?","intro":"64 synthetic uploads (titles, descriptions, transcripts, app listings) were labelled primary, mixed, general or mature from the public factor lists, with a made-for-kids setting each; hard cases include children on screen in content for parents and cartoons for adults. Prompts were tuned on 24 and frozen; 40 were held out. The frame check used 12 public-domain stills.","rows":[{"label":"Made for kids or not (held-out 40)","value":"39 of 40","detail":"the miss came back as a close call for a reviewer"},{"label":"Setting mismatches found","value":"15 of 15","detail":"no false mismatch"},{"label":"Audience, 4-way","value":"37 of 40","detail":"two mixed-audience uploads called primary"},{"label":"Audience-only audit","value":"39 of 40","detail":"a third of the calls"},{"label":"Frame check, picture only","value":"10 of 12","detail":"6 of 12 without it"}],"points":[{"heading":"Where it fails","text":"A stated child age range ('ages 6 and up') makes the model call a mixed-audience upload primary. A family road-trip vlog came back as a close call. Two turn-of-the-century engravings were not recognised as children's books from the picture alone."},{"heading":"What this does not show","text":"The uploads are short and synthetic, written and labelled by the same author. No real channel metadata, no independent labels and no legal judgement are measured."}],"source":"decosa-api docs/evals/kids-content-preflight.md, 2026-09-26"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":7200,"p95_ms":null,"runs":null,"receipts_per_run":15,"cost_per_run_usd":0.0045},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone into a clean directory, docker build of the api image (85 s), the api with a named volume on host networking against the running local vLLM (Qwen3.8-27B) and diarizer; then torn down.","notes":"The puppet sample gave made for kids, a high mismatch and the expected flags in 3.6 s (11 attested calls); the review record verified and failed once the designation was changed; the video sample was transcribed (6 lines, signed ASR receipt) and OCR'd in the container; the rehearsal bundle passed 13 of 13; no titles or transcript text in the logs. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route); production gets this tool when the branch merges.","Measured on 64 synthetic uploads written and labelled by the building agent; not on real channel metadata or with an independent reviewer's labels.","Mixed-audience uploads with a stated child age range are usually called primary; the made-for-kids designation is still right.","The audience probability comes from sampling and moves between runs; one sample trailer landed at 0.41 on one run and 0.93 on another.","The frame check was measured on 12 public-domain stills, not on real kids' video."],"receipt_coverage":"full"},"cost_per_run_usd":0.0045,"rehearsal_bundle":{"url":"/samples/kids-content-preflight.zip","checks":13,"bytes":4455},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Evidence lines per FTC factor (quotes), a typed audience choice with probabilities, a typed yes/no per factor (typed-judgment, vertical 24); with vision on, a description of four frames","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Speech to timed lines for a video sent without a transcript, with a speech receipt signed by the instance","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Tesseract OCR 5 (English)","role":"Reads on-screen text from six sampled frames of a video","license":"Apache-2.0","hf_repo":null},{"name":"decosa-api kids module (decosa_api/verticals/kids)","role":"The pre-flight: intake, quote checks, word-list cues, the setting check, the kids-directed flags, the batch audit, sign-off and the signed review record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/kids-content-preflight","page":"/tools/media/kids-content-preflight","json":"/use-cases/kids-content-preflight.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"hcc-evidence-file","num":"62","name":"HCC evidence file and RADV defence","status":"live","industries":["healthcare","finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Verdict accuracy, 3 classes","value":"189 / 200","unit":null,"n":200,"split":"test","note":"50 synthetic members; supported, not supported, insufficient."},{"name":"Supported vs not supported","value":"96 / 100","unit":null,"n":100,"split":"test","note":"Codes planted as supported or not supported."},{"name":"Unsupported or held codes called supported (false keep)","value":"0 / 161","unit":null,"n":161,"split":"test","note":null},{"name":"Supported codes called not supported (false delete)","value":"4 / 39","unit":null,"n":39,"split":"test","note":"3 are 'seropositive rheumatoid arthritis' read as not stating rheumatoid factor."},{"name":"Supported calls whose quotes include a planted evidence sentence","value":"35 / 35","unit":null,"n":35,"split":"test","note":null},{"name":"Quoted sentences that are planted evidence","value":"70 / 72","unit":null,"n":72,"split":"test","note":null},{"name":"Add suggestions in any output string or the evidence file","value":"0 (50 members)","unit":null,"n":50,"split":"test","note":null},{"name":"Verdict accuracy, 3 classes (dev)","value":"46 / 48","unit":null,"n":48,"split":"dev","note":"42 / 48 before the one prompt change made on dev."},{"name":"Verdict accuracy, scanned charts (document reader)","value":"187 / 200","unit":null,"n":200,"split":"test","note":"Same members as text in the same run: 192 / 200. Synthetic scans."},{"name":"Same verdict, scanned vs text","value":"195 / 200","unit":null,"n":200,"split":"test","note":"All 5 differences deleted or held. A later split found one false keep (fixed; 50 new members after: none)."}],"dataset":"Synthetic members from the vertical's own generator (invented people, providers and notes): 15 conditions that map to V28 payment HCCs, each code planted as supported, history, uncertain, less specific, absent, or evidence only in a record RADV does not accept. Dev 12 members (48 codes), test 50 members (200 codes), run once after the dev work was frozen.","held_out":true,"caveats":["Everything is synthetic, from sentence templates written by the same agent that wrote the checker and prompts: evidence the mechanisms work, not accuracy on real charts.","No certified-coder labels: a 200-chart coder-labelled set is the next step.","Scanned charts are read by the document reader first; its scans are clean synthetic pages (typed notes, no handwriting in the body), so real faxes will read worse. Text input is unchanged.","9 of the 11 test errors are one condition's wording ('seropositive' not read as 'with rheumatoid factor'); not fixed, since it showed up on test.","V28, payment year 2026 only."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/hcc-evidence-file"},"quality_evidence":[{"tier":"lite","label":"Lite · map and record checks only, no GPU","evidence":[{"metric":"V28 table against CMS's files","value":"8,019 codes, 115 HCC labels, 60 hierarchies; unit-tested against CMS's mapping for spot codes and the age edit on C50.911","source":"tests/test_hcc.py, 26 Sep 2026"},{"metric":"Record-check accuracy on real charts","value":"not measured yet","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Verdict accuracy, 3 classes (50 held-out synthetic members, 200 codes)","value":"189 / 200","source":"docs/evals/hcc-evidence-file.md, test split, 26 Sep 2026"},{"metric":"Supported vs not supported (codes planted as one or the other)","value":"96 / 100","source":"docs/evals/hcc-evidence-file.md, test split"},{"metric":"Codes that should be deleted or held but were kept","value":"0 / 161","source":"docs/evals/hcc-evidence-file.md, test split"},{"metric":"Supported calls whose quotes include a planted evidence sentence","value":"35 / 35","source":"docs/evals/hcc-evidence-file.md, test split"},{"metric":"Add suggestions or unsubmitted codes in the output","value":"0 in 50 members","source":"docs/evals/hcc-evidence-file.md, test split"}]},{"tier":"best","label":"Best · scanned charts too (adds the document reader)","evidence":[{"metric":"Verdict accuracy on scanned charts vs the same members as text (50 held-out synthetic members, 200 codes, same run)","value":"187 / 200 vs 192 / 200","source":"decosa-api docs/evals/document-reader.md, HCC test split, 26 Sep 2026"},{"metric":"Same verdict, scanned vs text","value":"195 / 200","source":"docs/evals/document-reader.md, test split"},{"metric":"Codes that should be deleted or held but were kept, on scans (test / fresh 1 / fresh 2)","value":"0 / 1 / 0 of about 160 each","source":"docs/evals/document-reader.md: fresh 1 found a coder's addendum read into the note; fixed, then fresh 2 (50 new members) had none"},{"metric":"Record header fields read right from the scans (date, type, provider, credential, signed)","value":"181 / 182 each","source":"docs/evals/document-reader.md, test split"},{"metric":"Quotes cited to a page and box of the scan","value":"315 / 315","source":"docs/evals/document-reader.md, test split"},{"metric":"After both reader fixes, 50 new members: scanned vs text","value":"189 / 200 vs 190 / 200","source":"docs/evals/document-reader.md, fresh2 split"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":23863,"p95_ms":null,"runs":null,"receipts_per_run":15,"cost_per_run_usd":0.003061},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_HCC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 15 of 15 checks in 4.8 s, the smoke module passed in 2.0 s with 9 attested receipts, a member not marked synthetic streamed a full review, and the tampered record failed to verify. Model-server startup itself not re-verified."},"known_limits":["Synthetic only on the hosted demo, and every number here comes from our own synthetic members; not run on real charts or against coder labels.","This page takes typed notes: dates, signatures and credentials sent as fields and text. Scanned charts are read only through the API (POST /hcc/read-chart with an API key) or self-hosted; the page has no upload for them yet.","V28 for payment year 2026 (2025 dates of service) only; no V24 or blended years.","'Not supported' means not supported by the notes sent; a coder may find another record.","It does not compute risk scores or payment amounts, and it does not run the RADV coversheet, attestation or member-identity checks."],"receipt_coverage":"full"},"cost_per_run_usd":0.003061,"rehearsal_bundle":{"url":"/samples/hcc-evidence-file.zip","checks":15,"bytes":3316},"models":[{"name":"decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)","role":"V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/hcc-evidence-file","page":"/clinics/hcc-evidence-file","json":"/use-cases/hcc-evidence-file.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"vex-triage","num":"63","name":"Signed VEX triage","status":"live","industries":["software","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"not_affected precision against Canonical's VEX","value":"154 / 159","unit":null,"n":159,"split":"test","note":"held-out ubuntu:jammy-20220421, run once with the code frozen"},{"name":"False not_affected on vendor-affected findings","value":"5 / 378","unit":null,"n":378,"split":"test","note":"all five supporting checks hold on the image (other platform, or a program from another package)"},{"name":"not_affected recall","value":"154 / 474","unit":null,"n":474,"split":"test","note":"Canonical's not_affected statements for packages in the image, added as probes"},{"name":"under_investigation rate","value":"106 / 852","unit":null,"n":852,"split":"test","note":"98 of the 106 are findings Canonical calls not_affected"},{"name":"Justification equal to Canonical's","value":"151 / 154","unit":null,"n":154,"split":"test","note":"when both say not_affected"},{"name":"Supporting checks that hold when re-done with other tools","value":"159 / 161","unit":null,"n":161,"split":"test","note":"readelf, find, grep, file, packaging.version; 0 fail, 2 could not be parsed"},{"name":"Rules only (no model): not_affected precision","value":"159 / 164","unit":null,"n":164,"split":"test","note":"same checks, no model call"},{"name":"not_affected precision, dev set","value":"19 / 24","unit":null,"n":24,"split":"dev","note":"ubuntu:focal-20210416, where prompts and code were tuned"}],"dataset":"Grype 0.119 findings on ubuntu:focal-20210416 (dev, 458 findings) and ubuntu:jammy-20220421 (test, 912), plus Canonical's not_affected statements for packages in each image as probe findings; labels from Canonical's OpenVEX feed, 26 Sep 2026.","held_out":true,"caveats":["One vendor's labels (Canonical) on two releases; Canonical writes about source packages and this tool about the image, which explains all five risky errors on the test set.","Probe findings are constructed, so precision depends on the mix; read the per-class counts.","The model adds little to the decisions: rules only reach the same precision. The building agent read the statements; no independent security engineer.","Distro packages only; language packages were not measured."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/vex-triage"},"quality_evidence":[{"tier":"lite","label":"Lite · rules only, no GPU","evidence":[{"metric":"not_affected precision against Canonical's VEX, held-out jammy set","value":"159/164","source":"decosa-api docs/evals/vex-triage.md, 2026-09-26 (rules-only baseline on the same run)"},{"metric":"false not_affected on vendor-affected findings, held-out","value":"5/378","source":"decosa-api docs/evals/vex-triage.md, 2026-09-26"},{"metric":"not_affected recall, held-out","value":"159/474","source":"decosa-api docs/evals/vex-triage.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"not_affected precision against Canonical's VEX, held-out jammy set (912 findings, run once)","value":"154/159","source":"decosa-api docs/evals/vex-triage.md, measured on our server 2026-09-26, gateway route; code frozen on the focal dev set"},{"metric":"false not_affected on vendor-affected findings (the risky error), held-out","value":"5/378; all five checks hold on the image (other platform, or a program from another package) but Canonical scores the source package","source":"decosa-api docs/evals/vex-triage.md, 2026-09-26"},{"metric":"not_affected recall / under_investigation rate, held-out","value":"154/474 / 106/852","source":"decosa-api docs/evals/vex-triage.md, 2026-09-26"},{"metric":"supporting checks re-done with readelf, find and grep","value":"159/161 hold, 0 fail (2 could not be parsed)","source":"decosa-api docs/evals/vex-triage.md, scripts/vex_evidence_audit.py, 2026-09-26"},{"metric":"labels from Red Hat, Debian or a security engineer","value":"not measured yet","source":null}]}],"benchmark":{"title":"Does it only say not affected when it can show why?","intro":"Grype findings on two public Ubuntu images, scored against Canonical's own OpenVEX for the same CVE and package. Canonical's not_affected statements for packages in the image were added as probes, shaped like an upstream-matching scanner's findings. Everything was tuned on ubuntu:focal-20210416; ubuntu:jammy-20220421 was run once with the code frozen.","rows":[{"label":"not_affected that Canonical agrees with","value":"154 of 159","detail":"held-out jammy set; dev 19 of 24"},{"label":"Vendor-affected findings wrongly marked not_affected","value":"5 of 378","detail":"all five checks hold on the image: 32-bit, PowerPC or Windows only, or a program from another package"},{"label":"Canonical's not_affected found","value":"154 of 474","detail":"the rest rest on a maintainer's judgment this tool never makes"},{"label":"nginx:1.20.0 findings closed with strong evidence","value":"168 of 550","detail":"rules mode; mostly libraries only unloaded modules pull in; no vendor labels for this image"}],"points":[{"heading":"Where it disagrees","text":"Its not_affected answers are about the image, and a distro's VEX is about the source package: an AES-OCB bug that only affects 32-bit x86 is not in the execute path of an amd64 image, but Canonical still ships the fix and calls the package affected. All five disagreements on the held-out set are of that kind."},{"heading":"What the model adds","text":"Little to the decisions: the rules alone reach the same precision (159 of 164). The model writes the impact statement with quotes and sends unclear cases to under investigation; the gates turned 101 of its unsupported not_affected proposals into under investigation."}],"source":"decosa-api docs/evals/vex-triage.md, 2026-09-26"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":14264,"p95_ms":null,"runs":null,"receipts_per_run":12,"cost_per_run_usd":0.014069},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing, DECOSA_VEX_OFFLINE=1; torn down after","notes":"The assembly prompt's smoke test passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): libwebp CVE-2023-4863 not_affected (not in execute path), CVE-2022-0778 affected, CVE-2009-4487 under investigation, all receipts attested, envelope and record verified, 2.0 s; the rehearsal bundle passed 10/10. The collector's --image path (crane) found that absolute symlinks were dropped when unpacking; fixed in b5a465e and re-checked (5,370 files). Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured against one vendor's VEX (Canonical) on two Ubuntu releases, with constructed probe findings; not against Red Hat or Debian data or a security engineer's review.","Recall of not_affected is about one in three: distro VEX often rests on maintainer judgment this tool does not make.","Language packages (PyPI, npm, Go) use the same OSV range check but were not measured against labels; there is no language-level reachability.","NVD allows 5 requests per 30 s without a key, so the hosted service fetches few new NVD records per run; findings then rest on OSV and the scanner's own data."],"receipt_coverage":"full"},"cost_per_run_usd":0.014069,"rehearsal_bundle":{"url":"/samples/vex-triage.zip","checks":10,"bytes":66935},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"One typed judgment per finding the rules do not settle: a VEX status with quoted evidence from the advisory lines and the checks, and the impact statement. Code gates every answer.","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa VEX checks and collector (code, no model)","role":"The collector (on your machine) and the checks in decosa-api: package in the SBOM, version against the fix and the OSV/NVD ranges, the Debian changelog, ELF loader chains, the advisory's programs and settings, the platform; the evidence gates; OpenVEX, CycloneDX VEX and DSSE signing.","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/vex-triage","page":"/tools/developer/vex-triage","json":"/use-cases/vex-triage.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"device-mdr-triage","num":"64","name":"Device complaint MDR triage","status":"live","industries":["healthcare","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Reportable complaints called not reportable","value":"6 of 96 (6.3%)","unit":null,"n":96,"split":"test","note":"The risky direction. Wilson 95% interval 2.9% to 13.0%. MAUDE 5 of 60, synthetic 1 of 36; 2 clear errors, 4 conservative filings"},{"name":"Not-reportable complaints called reportable","value":"0 of 58","unit":null,"n":58,"split":"test","note":"Complaints written and labelled by a separate agent from written rules"},{"name":"Not-reportable complaints cleared","value":"53 of 58","unit":null,"n":58,"split":"test","note":"The other 5 went to needs investigation (3 were another maker's device with an injury)"},{"name":"MAUDE reported events called reportable","value":"43 of 60","unit":null,"n":60,"split":"test","note":"12 needs investigation, 5 not reportable; labels are the filers' decisions"},{"name":"Synthetic reportable complaints called reportable","value":"34 of 36","unit":null,"n":36,"split":"test","note":"1 needs investigation, 1 not reportable"},{"name":"Accuracy on decisive answers","value":"130 of 136 (95.6%)","unit":null,"n":136,"split":"test","note":"Reportable or not reportable answers only"},{"name":"Clock deadline right","value":"36 of 36","unit":null,"n":36,"split":"test","note":"Synthetic clock and positive cases; gold dates hand-computed and script-checked by the labeller. MAUDE date checks 60 of 60"},{"name":"Narrative sentences kept by the grounding check","value":"43 of 43","unit":null,"n":43,"split":"test","note":"12 reportable test complaints; the check's own verdicts, no separate reader"}],"dataset":"154 test and 34 dev complaints: public openFDA MAUDE event narratives from 2025 (CC0; shortened, maker names replaced; labelled reportable because they were reported) and synthetic complaints about fictional devices written and labelled by a separate agent from rules drawn from 21 CFR 803.","held_out":true,"caveats":["MAUDE labels are the filers' decisions, which lean towards reporting; some misses are conservative filings an RA specialist could have closed.","The synthetic labels are one labeller's (an agent working from written rules), not an RA specialist's; no real complaint files were used.","One prompt change was made on the dev set before the test run; the test set was run once. One test item hit a date-range bug, fixed and rerun on its own.","The narrative check was measured on 12 complaints with the check's own verdicts only; the trend grouping on the demo trend only.","US FDA rules only; latency depends on the shared gateway's load."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/device-mdr-triage"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"false not-reportable rate and clock accuracy","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"test set (154 complaints, run once): false \"not reportable\"","value":"6 of 96 gold-reportable (6.3%); MAUDE 5 of 60, synthetic 1 of 36","source":"decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26, gateway route; prompts frozen after one change on a 34-complaint dev set"},{"metric":"false \"reportable\" / not-reportable complaints cleared","value":"0 of 58 / 53 of 58 (5 went to needs investigation)","source":"decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; complaints written and labelled by a separate agent from written rules"},{"metric":"reportable caught (MAUDE / synthetic)","value":"43 of 60 / 34 of 36; the rest went to needs investigation except the 6 above","source":"decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26"},{"metric":"clock deadline right","value":"36 of 36 synthetic cases (earlier employee awareness, FDA 5-day requests, remedial action, user facilities, holidays); 60 of 60 MAUDE date checks","source":"decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26; gold dates hand-computed and script-checked by the labeller"},{"metric":"narrative sentences kept by the grounding check","value":"43 of 43 on 12 reportable test complaints (its own verdicts; no separate reader)","source":"decosa-api docs/evals/device-mdr-triage.md, measured on our server 2026-09-26"},{"metric":"real complaint files labelled by an RA specialist","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"false not-reportable rate and clock accuracy","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"false not-reportable rate and clock accuracy","value":"not measured yet","source":null}]}],"benchmark":{"title":"How often does it clear a complaint that should be reported?","intro":"154 complaints run once: 60 public MAUDE event narratives (reported to FDA, so labelled reportable by the filer's decision) and 94 synthetic complaints written and labelled by a separate agent from written rules, without seeing the prompts. Prompts were changed once on a 34-complaint dev set, then frozen.","rows":[{"label":"Reportable complaints called not reportable","value":"6 of 96","detail":"6.3%; 2 clear errors, 4 conservative MAUDE filings"},{"label":"Not-reportable complaints called reportable","value":"0 of 58","detail":"53 cleared, 5 sent to needs investigation"},{"label":"Clock deadlines right","value":"36 of 36","detail":"earlier employee awareness, 5-day and 10-work-day clocks, holidays"},{"label":"Cost per triage with the narrative","value":"about $0.004","detail":"11 model calls, 9,900 tokens on the catheter demo, gateway list price"}],"points":[{"heading":"Where it fails","text":"Judging whether a malfunction with no injury would be likely to cause serious harm if it recurred: it cleared an infusion pump whose anti-free-flow clamp failed with no harm. It also missed that a planned hip revision is surgery. Terse MAUDE narratives often go to needs investigation (12 of 60)."},{"heading":"What it does not show","text":"MAUDE labels are the filers' decisions, which lean towards reporting, and the synthetic labels are one labeller's, not an RA specialist's. Real complaint files are longer and messier. Measure it on your own closed complaints before relying on it."}],"source":"decosa-api docs/evals/device-mdr-triage.md, 26 Sep 2026"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":11900,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0019},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): clocks 2026-10-02 and 2026-09-10; rep-told-earlier reportable, clock from 2026-09-02, due 2026-10-02, with a narrative in 4.2 s (10 attested receipts); cpap-lid-cosmetic not reportable; meter-reads-high needs investigation; triage and decision records verified; the rehearsal bundle passed 16/16. Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges. p50 is the eval's test median without the narrative; receipts and cost are from the smoke run (5 calls).","False not-reportable rate 6 of 96 on the test set: it can clear a reportable complaint, so a qualified person reviews every suggestion.","Measured on public MAUDE narratives (filer labels) and synthetic complaints (one labeller's labels); not on real complaint files or with an RA specialist's labels.","US FDA rules only (21 CFR 803, 820.35). No event problem codes, eMDR filing, remedial-action decisions or EU MDR vigilance.","The trend grouping and the two-year presumption were checked on the 7-complaint demo trend only."],"receipt_coverage":"full"},"cost_per_run_usd":0.0019,"rehearsal_bundle":{"url":"/samples/device-mdr-triage.zip","checks":16,"bytes":4622},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the complaint (dates an employee was told, the event date, the device problem), answers the 803.50(a) questions one at a time (outcome, caused or contributed, malfunction, likely if it recurred), drafts the 3500A event description, judges every narrative sentence (the grounding judge), and groups complaints by failure mode for the trend view","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/device-mdr-triage","page":"/tools/life-sciences/device-mdr-triage","json":"/use-cases/device-mdr-triage.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"cmmc-evidence-map","num":"65","name":"CMMC / NIST 800-171 evidence map","status":"live","industries":["compliance-trust","public-sector"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Objective status, 3 classes","value":"237 / 286","unit":null,"n":286,"split":"test","note":"10 synthetic companies, 80 requirements; present, partial, missing."},{"name":"Partial or missing objectives called present (false present)","value":"0 / 121","unit":null,"n":121,"split":"test","note":null},{"name":"Requirements called fully evidenced that are not","value":"0 / 59","unit":null,"n":59,"split":"test","note":null},{"name":"Present objectives called present (recall)","value":"119 / 165","unit":null,"n":165,"split":"test","note":"It errs down: 47 of 49 errors call an objective less evidenced than its label."},{"name":"Quotes verbatim from the input","value":"446 / 446","unit":null,"n":446,"split":"test","note":null},{"name":"Present or partial calls quoting a planted evidence line","value":"204 / 209","unit":null,"n":209,"split":"test","note":null},{"name":"Requirement roll-up accuracy","value":"64 / 80","unit":null,"n":80,"split":"test","note":null},{"name":"Objective status, self-hosted direct route","value":"235 / 286","unit":null,"n":286,"split":"test","note":"Log-probabilities instead of votes; false present 0 / 121. One rule change made on the self-hosted dev set first."},{"name":"Objective status, 3 classes (dev)","value":"238 / 299","unit":null,"n":299,"split":"dev","note":"215 / 299 before the dev changes; false present on dev 3 / 102."}],"dataset":"Synthetic small defence contractors from the vertical's own generator (invented companies, people and systems): 8 requirements each from a pool of 14 NIST SP 800-171 Rev 2 requirements (51 objectives), each with at most one planted gap (evidence missing, planned, stale, out of scope, draft, shown in part, contradicted, one line dropped). Dev 10 companies (299 objectives, one phrasing), test 10 companies (286 objectives, every evidence line worded differently), run once after the dev work was frozen.","held_out":true,"caveats":["Everything is synthetic, from templates written by the same agent that wrote the prompts and rules: evidence the mechanisms work, not accuracy on real SSPs.","No RPO or assessor labels; 14 of the 110 requirements are exercised.","Labels count only a requirement's own artefacts; the model sometimes uses evidence filed under another requirement.","Screenshot reading is shown on one bundled image, not measured.","Measured on a gateway shared with other workloads, so the latencies are high and vary."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/cmmc-evidence-map"},"quality_evidence":[{"tier":"lite","label":"Lite · catalog and artefact checks only, no GPU","evidence":[{"metric":"Catalog against NIST's CSVs and 32 CFR 170","value":"110 requirements, 320 objectives; 42 five-point, 14 three-point, 2 variable, 52 one-point; the six never-POA&M requirements; unit-tested","source":"tests/test_cmmc.py, 26 Sep 2026"},{"metric":"Artefact-check accuracy on real evidence","value":"not measured yet","source":"not measured yet"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Objective status, 3 classes (10 held-out synthetic companies, 286 objectives)","value":"237 / 286","source":"docs/evals/cmmc-evidence-map.md, test split, 26 Sep 2026"},{"metric":"Partial or missing objectives called present (false present)","value":"0 / 121","source":"docs/evals/cmmc-evidence-map.md, test split"},{"metric":"Requirements called fully evidenced that are not","value":"0 / 59","source":"docs/evals/cmmc-evidence-map.md, test split"},{"metric":"Present objectives called present (recall)","value":"119 / 165","source":"docs/evals/cmmc-evidence-map.md, test split"},{"metric":"Quotes verbatim from the input","value":"446 / 446","source":"docs/evals/cmmc-evidence-map.md, test split"},{"metric":"Objective status, self-hosted direct route (same 286 test objectives)","value":"235 / 286, false present 0 / 121","source":"docs/evals/cmmc-evidence-map.md, self-hosted test run"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":149546,"p95_ms":null,"runs":null,"receipts_per_run":33,"cost_per_run_usd":0.0078},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_CMMC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 21 of 21 checks in 19.8 s, the smoke module passed in 5.3 s with 18 attested receipts, unmarked and CUI-marked sets were accepted as real-data mode should, and the 10 held-out test companies scored 235 of 286 with no false present. Model-server startup itself not re-verified."},"known_limits":["Synthetic only on the hosted demo, and every number here comes from our own synthetic companies over 14 of the 110 requirements; not run on real SSPs or against an assessor's labels.","Reads text, config exports and single screenshots; the hosted demo reads only the bundled sample screenshot. Scanned PDFs and Word binders are not read.","It errs down: recall of 'present' is 72% on the test set, so some evidenced objectives come back partial.","'Missing' means not in what was sent. It does not interview, test or examine systems as an assessor does.","No SPRS score. A self-assessed estimate appears only when all 110 requirements are mapped in one run (self-hosted), labelled as an estimate.","NIST SP 800-171 Rev 2 and CMMC Level 2 only; no Level 1, Level 3 or Rev 3."],"receipt_coverage":"full"},"cost_per_run_usd":0.0078,"rehearsal_bundle":{"url":"/samples/cmmc-evidence-map.zip","checks":21,"bytes":67272},"models":[{"name":"decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)","role":"Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Screenshot transcription, lines per objective, the typed judgment per objective, and the grounding judge","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/cmmc-evidence-map","page":"/tools/finance/cmmc-evidence-map","json":"/use-cases/cmmc-evidence-map.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"eu-trial-lay-summary","num":"66","name":"EU trial lay summary with number grounding","status":"live","industries":["healthcare","science-research"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted errors caught, drafts with citations (code checks)","value":"555 / 577 (96.2%)","unit":null,"n":577,"split":"test","note":"12 held-out ClinicalTrials.gov trials; numbers, percentages, dates, swapped groups, flipped comparisons, dropped hedges, softened or denied side effects."},{"name":"Correct sentences flagged (false flags), code checks","value":"9 / 812 (1.1%)","unit":null,"n":812,"split":"test","note":"Adjudicated by the author."},{"name":"Planted errors caught after fixes, fresh trials","value":"212 / 217 (97.7%)","unit":null,"n":217,"split":"heldout","note":"5 trials kept aside while the test split's misses were fixed; 16 of 347 correct sentences flagged there, from one trial's shortened group names (fixed afterwards)."},{"name":"Planted errors caught, citations stripped","value":"315 / 577 (54.6%)","unit":null,"n":577,"split":"test","note":null},{"name":"Wrong numbers in the model's first drafts","value":"5 / 642 (0.8%)","unit":null,"n":642,"split":"test","note":"None left after the check and one repair pass (0 / 655)."},{"name":"Planted wording errors caught by the grounding judge","value":"29 / 45 (64%)","unit":null,"n":null,"split":"test","note":null},{"name":"Flesch-Kincaid grade, median (range)","value":"6.1 (4.4-8.0)","unit":null,"n":12,"split":"test","note":null},{"name":"Planted errors caught, dev","value":"130 / 132","unit":null,"n":132,"split":"dev","note":"The 3 sample trials, used while writing the prompts and the checker."},{"name":"Numbers in Member-State versions traced to the results, 6 languages (frozen)","value":"4,002 / 4,205 (95.2%)","unit":null,"n":4205,"split":"test","note":"The 12 test trials' English drafts translated into de, fr, es, it, nl, pl. Most untraced numbers were the checker misreading other languages (fixed afterwards: 4,014 / 4,077, no longer held out)."},{"name":"Planted number errors in the translations caught","value":"2,426 / 2,432 (99.8%)","unit":null,"n":2432,"split":"test","note":"One digit changed in a sentence both checks had passed; 917 / 920 on the 5 test2 trials. Real errors the check found in the model's translations: 6 Polish sentences that dropped a '0 out of' count (test2) and one Spanish '30,000' left in English format."}],"dataset":"20 real phase 3 trials with results on ClinicalTrials.gov (19 sponsors, fetched 26 Sep 2026): dev 3, test 12 (run once, frozen), test2 5 (kept fresh for the fixes the test split led to). Drafts by our model; errors planted in code, one per sentence. Member-State versions: the test and test2 drafts translated into six languages on 27 Sep 2026.","held_out":true,"caveats":["The drafts and the planted errors are ours; no summaries written by people and no published lay summaries were checked.","False flags and the 10-summary review were judged by the agent that built the checker; no independent reviewer and no lay-reader test.","After the test run its misses were fixed; the test numbers after the fixes (576 / 577) are no longer held out, and the fresh test2 run (212 / 217) checks the fixes.","ClinicalTrials.gov records only; CTIS and EudraCT tables were not read.","Coverage says which Annex V parts are present, not whether they are adequate; it never says a summary complies.","Member-State versions are machine translations with numbers checked in code; the optional meaning check caught all 6 known Polish \"0 out of\" drops and real errors such as \"Vehicle Cream\" as a cream for cars, but also raises false alarms (9 of 14 errors on a 240-sentence sample). Grammar is not checked; no native speaker has read them.","The fifteen EU languages added on 27 Sep (Qwen3.8-27B route) were run on the 5 test2 trials in five of them (fi, el, hu, ro, ga): 1,604 of 1,624 numbers traced to the results, 1 wrong (Irish); real errors found were a dropped Finnish count and two wrong Irish month names. No native speaker has read any version."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/eu-trial-lay-summary"},"quality_evidence":[{"tier":"lite","label":"Lite · check a draft in code, no GPU","evidence":[{"metric":"Planted errors caught in drafts with citations (12 held-out trials)","value":"555 / 577 (96.2%)","source":"docs/evals/eu-trial-lay-summary.md, test split, 26 Sep 2026"},{"metric":"Correct sentences flagged (false flags, same drafts)","value":"9 / 812 (1.1%)","source":"docs/evals/eu-trial-lay-summary.md, test split, adjudicated by the author"},{"metric":"Planted errors caught with the citations stripped","value":"315 / 577 (54.6%)","source":"docs/evals/eu-trial-lay-summary.md, test split"}]},{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Wrong numbers in the model's first drafts (12 held-out trials)","value":"5 / 642 (0.8%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Wrong numbers left after the check and one repair","value":"0 / 655","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Planted errors caught by the code checks (12 held-out trials)","value":"555 / 577 (96.2%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Planted wording errors caught by the grounding judge","value":"29 / 45 (64%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Flesch-Kincaid grade of the drafts: median (range)","value":"6.1 (4.4-8.0)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Numbers in Member-State versions traced to the results cells (12 held-out trials, 6 languages, frozen)","value":"4,002 / 4,205 (95.2%)","source":"decosa-api docs/evals/language-pack.md, 27 Sep 2026"},{"metric":"Planted number errors in the translations caught (12 trials, 6 languages)","value":"2,426 / 2,432 (99.8%)","source":"decosa-api docs/evals/language-pack.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":8248,"p95_ms":null,"runs":null,"receipts_per_run":3,"cost_per_run_usd":0.001406},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_LAYSUMMARY_FETCH=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 9 of 9 checks in 0.8 s, the smoke module passed in 0.9 s with 3 attested receipts, a full draft of the ruxolitinib sample took 18.4 s (59 attested receipts, 49 of 49 numbers traced, grade 5.7) and its record verified, and an NCT number was refused with fetching off. Torn down after. Model-server startup itself not re-verified."},"known_limits":["Drafts in English; Member-State versions are machine translations with their numbers checked in code, not their wording.","Reads ClinicalTrials.gov records; CTIS and EudraCT results tables are not read.","Without citations in a draft, about half of the planted wrong numbers were missed (315 of 577 caught): cite the cells or turn the grounding judge on.","The grounding judge can be wrong and is not fully repeatable on the shared gateway; one correct sentence was once called contradicted.","Coverage says which Annex V parts are present, never that a summary meets Annex V."],"receipt_coverage":"full"},"cost_per_run_usd":0.001406,"rehearsal_bundle":{"url":"/samples/eu-trial-lay-summary.zip","checks":9,"bytes":17111},"models":[{"name":"decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)","role":"Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sources","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Hy-MT2-7B (the language-pack block)","role":"Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Qwen3.8-27B (the language-pack block's route for these languages)","role":"Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)","license":"Apache-2.0","hf_repo":"Qwen/Qwen3.8-27B"}],"licence":"permissive","links":{"metrics":"/metrics/eu-trial-lay-summary","page":"/tools/life-sciences/eu-trial-lay-summary","json":"/use-cases/eu-trial-lay-summary.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"reg-e-dispute-file","num":"67","name":"Reg E dispute investigation file","status":"live","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Clocks agreeing with a second implementation","value":"2,763 / 2,763","unit":null,"n":2763,"split":"synthetic","note":"600 random disputes, 2026 to 2030; the oracle was written by the same author"},{"name":"Missing notice elements flagged (recall)","value":"0.956 (43/45)","unit":null,"n":45,"split":"test","note":"40 synthetic notices, run once after the prompts were frozen"},{"name":"Missing-element flags that were right (precision)","value":"0.977 (43/44)","unit":null,"n":44,"split":"test","note":null},{"name":"Notices with all five elements right","value":"37 / 40","unit":null,"n":40,"split":"test","note":"name, account, why, date, amount"},{"name":"Letter explanation specific or not","value":"20 / 20","unit":null,"n":20,"split":"test","note":"templated letters"},{"name":"Right-to-documents sentence found or missing","value":"20 / 20","unit":null,"n":20,"split":"test","note":"includes a paraphrase no keyword rule catches"},{"name":"Debit notice complete or not","value":"20 / 20","unit":null,"n":20,"split":"test","note":null},{"name":"Automated decisions","value":"0 / 9","unit":null,"n":9,"split":"synthetic","note":"5 samples and 4 injection attempts on the real model"}],"dataset":"56 synthetic notices (16 dev, 40 test) from 12 dispute scenarios in four channels; 27 templated results letters (7 dev, 20 test); 600 random disputes for the clocks; 5 demo files and 4 injection attempts.","held_out":true,"caveats":["The same author wrote the cases, the gold labels and the prompts.","Synthetic and templated only; no real bank notices or letters.","Letters come from a small set of variants: 20/20 shows the check works on clear cases, not on your templates.","Who made the transfer is not in the headline: raw agreement 24/40, with label errors on our side; it only picks a coverage note.","The clock oracle checks the code, not the reading of the rule."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/reg-e-dispute-file"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"notice elements and letter checks","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"clocks: due dates and statuses against a second implementation (600 random disputes, 2026 to 2030, 2,763 clocks)","value":"2,763/2,763; the oracle was written by the same author, so this checks the code, not the reading of the rule","source":"decosa-api docs/evals/reg-e-dispute-file.md, 2026-09-26"},{"metric":"notice elements, test split (40 synthetic notices, run once): 'absent' flags","value":"precision 0.977 (43/44), recall 0.956 (43/45); all five elements right on 37/40","source":"decosa-api docs/evals/reg-e-dispute-file.md, measured on our server 2026-09-26, gateway route; prompts frozen on a 16-notice dev set"},{"metric":"denial-letter completeness, test split (20 templated letters)","value":"explanation 20/20, right to documents 20/20, debit notice 20/20; small templated set","source":"decosa-api docs/evals/reg-e-dispute-file.md, measured on our server 2026-09-26"},{"metric":"automated decisions (5 samples and 4 injection attempts on the real model, plus unit tests)","value":"0 of 9","source":"decosa-api docs/evals/reg-e-dispute-file.md, 2026-09-26"},{"metric":"real, redacted bank notices and letters","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"notice elements and letter checks","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"notice elements and letter checks","value":"not measured yet","source":null}]}],"benchmark":{"title":"Are the clocks right, and does it catch the letter examiners cite?","intro":"The clocks were checked against a second implementation on 600 random disputes. The notice and letter checks ran on synthetic notices and letters with planted gaps: prompts were written on a dev split, then a test split was run once.","rows":[{"label":"Clocks agreeing with the oracle","value":"2,763 of 2,763","detail":"600 disputes, 2026 to 2030, Federal Reserve holidays"},{"label":"Missing notice elements flagged","value":"43 of 45","detail":"test split; 1 false flag in 44"},{"label":"Generic 'no error' letters and missing right-to-documents sentences caught","value":"20 of 20 letters right on all three items","detail":"test split, templated"},{"label":"Automated decisions","value":"0 of 9","detail":"real model, including 4 injection attempts"}],"points":[{"heading":"Where it fails","text":"On notices, a detail that only hints at an account (\"my prepaid card\" in the provider's own chat) was read as identifying it, and a transfer confirmation number was not. Who made the transfer is a hint only: raw agreement 24 of 40, several of our labels were wrong, and it only picks which coverage note the investigator sees."},{"heading":"What it does not show","text":"Every notice and letter is synthetic and templated, written by the same author as the prompts. The CFPB complaint database's narratives were the planned real-world input, but its public API no longer returns them. Run your own notices and letter templates through the rehearsal bundle before relying on it."}],"source":"decosa-api docs/evals/reg-e-dispute-file.md, 26 Sep 2026"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":37750,"p95_ms":null,"runs":null,"receipts_per_run":8,"cost_per_run_usd":0.00568},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): clocks as expected, cnp-generic-denial gaps with cannot_close in 18.6 s (8 attested receipts), p2p-scam-consumer-sent gaps with the consumer_sent coverage question, p2p-takeover-open open with no determination, record verified; the rehearsal bundle passed 14/14. Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.","Measured on synthetic, templated notices and letters written by the building agent; not on a bank's real notices, notes or letter templates.","Federal rules only: no Regulation Z, card-network rules or state law. Business days default to the Federal Reserve Banks' calendar.","The notice, notes and letter are read by a model and can be wrong in both directions; the clocks depend on the dates you send (statement dates in the CSV)."],"receipt_coverage":"full"},"cost_per_run_usd":0.00568,"rehearsal_bundle":{"url":"/samples/reg-e-dispute-file.zip","checks":14,"bytes":5984},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the consumer's notice (the required elements, the error asserted, who made the transfer, the transfers named), the investigator's notes (evidence reviewed, waiting for paperwork, carelessness cited) and the results letter (explanation specific or generic, right to documents, debit notice), and judges each letter sentence against the file (the grounding judge)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/reg-e-dispute-file","page":"/tools/finance/reg-e-dispute-file","json":"/use-cases/reg-e-dispute-file.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"virtual-staging","num":"68","name":"Disclosed virtual staging","status":"live","industries":["creative-media","sales-marketing"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted edits caught (check)","value":"38 / 40","unit":null,"n":40,"split":"test","note":"Window removed 10/10, wall recoloured 10/10, view replaced 10/10, crack or damage hidden 8/10."},{"name":"False alarms on clean stagings","value":"0 / 10","unit":null,"n":10,"split":"test","note":null},{"name":"Planted edits caught, pixel check alone","value":"38 / 40","unit":null,"n":40,"split":"test","note":null},{"name":"Planted edits caught, model alone","value":"29 / 40","unit":null,"n":40,"split":"test","note":null},{"name":"Label read back after 1024 px and JPEG q70","value":"30 / 30","unit":null,"n":30,"split":"test","note":null},{"name":"First render passes the check","value":"18 of 32 (56%)","unit":null,"n":32,"split":"synthetic","note":"8 rooms x 4 seeds; 0 of 4 on a tiled floor with a border"}],"dataset":"8 CC0 empty-room photos (Poly Haven panoramas cut to views, Wikimedia Commons), split by room: 3 dev, 5 test. 17 clean stagings by this service labelled by eye; 62 planted Pillow edits (window removed, wall recoloured, view replaced, crack or damage hidden).","held_out":true,"caveats":["The same author wrote the planted edits and the checker and labelled the clean stagings; some edits are crude.","Small: 10 clean and 40 planted test cases from 5 rooms; no images staged by other tools.","Thresholds were tuned on dev in three rounds; the floor-pattern rule was set after seeing renders of a test room.","One redrawn floor patch in a staging that looked otherwise clean passed the check (0 of 1).","Not legal advice or a compliance determination."],"date":"2026-09-26","doc_url":"https://decosa.ai/metrics/evals/virtual-staging"},"quality_evidence":[{"tier":"lite","label":"Lite · check only (any staging tool)","evidence":[{"metric":"held-out test: planted edits caught (window removed, wall recoloured, view replaced, crack hidden) / false alarms on clean stagings","value":"38 of 40 / 0 of 10","source":"decosa-api docs/evals/virtual-staging.md, measured on our server 2026-09-26, gateway route"},{"metric":"burned-in label read back by OCR after resizing to 1024 px and JPEG quality 70","value":"30 of 30","source":"decosa-api docs/evals/virtual-staging.md, 2026-09-26"}]},{"tier":"standard","label":"Standard · stage and check (hosted demo)","evidence":[{"metric":"first render passes the check, 8 rooms x 4 seeds","value":"18 of 32 (0 of 4 on a tiled floor with a border, 0 of 4 on a cluttered room)","source":"decosa-api docs/evals/virtual-staging.md, 2026-09-26"},{"metric":"passing stagings that look furniture-only to a person (the building agent)","value":"17 of 18","source":"decosa-api docs/evals/virtual-staging.md, 2026-09-26"}]}],"benchmark":{"title":"Does the check catch a doctored listing photo?","intro":"Eight CC0 empty rooms were staged by this service; each honest staging was then doctored four ways with ordinary photo edits. Thresholds were tuned on three rooms and frozen; five rooms were held out.","rows":[{"label":"Planted edits caught (held-out 40)","value":"38 of 40","detail":"both misses are hairline cracks on a marble-effect wall"},{"label":"Honest stagings flagged","value":"0 of 10","detail":null},{"label":"Window removed / view replaced","value":"10 of 10 / 10 of 10","detail":null},{"label":"Wall recoloured / crack or damage hidden","value":"10 of 10 / 8 of 10","detail":null},{"label":"Label read back after resizing and JPEG q70","value":"30 of 30","detail":null}],"points":[{"heading":"Where it fails","text":"Hairline cracks on a patterned wall; a patch of floor redrawn around a chair on a floor with little pattern; and our own renderer on a tiled floor with a border, which it redraws (the check refuses those, so the room is not disclosed)."},{"heading":"What this does not show","text":"Images staged by other tools, careful retouching, independent labels and legal judgement are not measured. The planted edits are ordinary Pillow edits written by the same author as the checker."}],"source":"decosa-api docs/evals/virtual-staging.md, 2026-09-26"},"verification":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":15200,"p95_ms":null,"runs":null,"receipts_per_run":3,"cost_per_run_usd":0.0021},"selfhost":{"date":"2026-09-26","result":"pass","method":"Fresh clone into a clean directory, docker build of the api image (31 s), the api on host networking with named volumes against the running local vLLM (Qwen3.8-27B with vision) and ComfyUI; then torn down.","notes":"The rehearsal bundle passed 13 of 13 (doctored view refused, honest staging labelled, read back, credentialed and verified, record tamper caught, power-lines refused); a Berlin staging rendered and passed in 12.8 s with attested receipts. When ComfyUI runs outside the api container, set DECOSA_STUDIO_COMFY_INPUT_DIR or delete its input copies yourself."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route); production gets this tool when the branch merges.","Measured on 8 CC0 rooms with planted Pillow edits and our own renders, labelled by the building agent; images staged by other tools were not tested.","The renderer redraws patterned floors (a tile border) around furniture; the check then refuses to label it, so such rooms often end not disclosed.","Hairline cracks on patterned walls can be missed (2 of 3 on a marble-effect wall).","The C2PA credential uses a development certificate; public validators show it as untrusted."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0021,"rehearsal_bundle":{"url":"/samples/virtual-staging.zip","checks":13,"bytes":591408},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Maps the original's windows, doors, fixtures, damage and open floor; boxes each piece of furniture on the render; compares original and staged side by side; the typed judgment on instructions (typed-judgment, vertical 24)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Wan2.2-VACE-Fun-A14B","role":"Draws furniture inside the zone (the open floor, with windows, doors, radiators, built-ins and damage cut out): one inpainted still per attempt, VACE control image plus mask","license":"Apache-2.0","hf_repo":"alibaba-pai/Wan2.2-VACE-Fun-A14B"},{"name":"Wan2.2-Lightning T2V 4-step LoRAs","role":"4-step distillation LoRAs for the render","license":"Apache-2.0","hf_repo":"lightx2v/Wan2.2-Lightning"},{"name":"decosa-api staging module (decosa_api/verticals/staging)","role":"The zone, the furniture-only composite, the pixel check (added, removed, recoloured regions; floor pattern; shadows; what each touches), the verdict, the burned-in label and its OCR read-back, the C2PA credential, the public original page and the signed pair record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/virtual-staging","page":"/tools/operations/virtual-staging","json":"/use-cases/virtual-staging.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"ma-dd-redflags","num":"69","name":"M&A due-diligence red flags","status":"live","industries":["legal","finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted red flags caught","value":"20 / 20","unit":null,"n":20,"split":"test","note":"two held-out invented rooms (Osprey 11, Heron 9), wording set B, run once with prompts frozen"},{"name":"Flags raised in a clean room of near misses","value":"0","unit":null,"n":null,"split":"test","note":"held-out room Wren"},{"name":"Quotes found word for word at the cited byte span","value":"22 / 22","unit":null,"n":22,"split":"test","note":null},{"name":"Planted flags caught in 40-document rooms (48 public contracts added)","value":"20 / 20","unit":null,"n":20,"split":"test","note":"direct route; 0 flags raised in the added contracts"},{"name":"CUAD anti-assignment clauses found","value":"9 / 12","unit":null,"n":12,"split":"heldout","note":"28 public CUAD contracts, precision 0.75; BM25 only: 11 / 12"},{"name":"CUAD exclusivity and non-compete clauses found","value":"3 / 7","unit":null,"n":7,"split":"heldout","note":"precision 3/3; BM25 only: 4 / 7"},{"name":"Planted red flags caught, dev","value":"11 / 11","unit":null,"n":11,"split":"dev","note":"Kestrel; clean dev room Linnet: 0 flags"}],"dataset":"Five invented data rooms of 16 documents (dev Kestrel and clean Linnet; held-out Osprey, Heron and clean Wren, worded from a second clause set written before any model run); the held-out rooms padded with 48 CUAD v1 contracts (CC BY 4.0); 28 further CUAD contracts scored against CUAD's expert labels.","held_out":true,"caveats":["Same author wrote the rooms, the gold and the prompts; the rooms are short and tidy compared with real data rooms.","The invented rooms are at the ceiling: BM25 search and reading the whole room also found every planted flag, so they do not separate retrieval methods.","On public CUAD contracts BM25 did as well or better than the reranked hybrid; small n (28 contracts).","No real data room and no lawyer's review of the memos.","Stress and CUAD runs used the direct route to the same model because the gateway was down."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/ma-dd-redflags"},"quality_evidence":[{"tier":"lite","label":"Lite · BM25 retrieval, no retrieval service","evidence":[{"metric":"Planted flags caught, held-out rooms, BM25 only","value":"20 / 20, 0 in the clean room","source":"decosa-api docs/evals/ma-dd-redflags.md, test split (direct route), 27 Sep 2026"},{"metric":"CUAD public contracts: anti-assignment / exclusivity found, BM25 only","value":"11/12 (precision 0.917) / 4/7","source":"decosa-api docs/evals/ma-dd-redflags.md, CUAD check"}]},{"tier":"standard","label":"Standard · retrieval service + the model (hosted demo)","evidence":[{"metric":"Planted flags caught, held-out rooms (Osprey, Heron)","value":"20 / 20","source":"decosa-api docs/evals/ma-dd-redflags.md, test split, 27 Sep 2026"},{"metric":"Flags raised in the clean held-out room (Wren)","value":"0","source":"decosa-api docs/evals/ma-dd-redflags.md, test split"},{"metric":"Quotes found word for word at the cited byte span","value":"22 / 22 (held-out memo flags)","source":"decosa-api docs/evals/ma-dd-redflags.md, test split"},{"metric":"Planted flags caught with 48 public contracts added as distractors (40-document rooms)","value":"20 / 20, 0 flags in the added contracts","source":"decosa-api docs/evals/ma-dd-redflags.md, stress split (direct route)"},{"metric":"CUAD public contracts: anti-assignment / exclusivity / change of control found","value":"9/12 (precision 0.75) / 3/7 / 2/2","source":"decosa-api docs/evals/ma-dd-redflags.md, CUAD check, 28 contracts"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":92855,"p95_ms":null,"runs":null,"receipts_per_run":3,"cost_per_run_usd":0.002531},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running retrieval service, local signing; torn down after","notes":"The rehearsal bundle passed 13/13 in 101 s (Kestrel with its planted flags, the record verified and failed when changed, the clean Linnet room with no flags; every receipt attested). The retrieval service was not built from the compose file here: the running one was used."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.","Measured on five small invented rooms written by the same author as the prompts, and 28 public contracts; not on a real data room or against a lawyer's issues list.","Ten categories only; tax, employment, data protection, environmental and sanctions issues are not looked for.","Text only: scanned documents must go through the document reader block first.","'No flag found' means nothing turned up in the passages the search returned for that category; a run where model calls failed is marked incomplete."],"receipt_coverage":"full"},"cost_per_run_usd":0.002531,"rehearsal_bundle":{"url":"/samples/ma-dd-redflags.zip","checks":13,"bytes":16440},"models":[{"name":"decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)","role":"Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each finding","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","role":"Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per search","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Reranker-4B"}],"licence":"permissive","links":{"metrics":"/metrics/ma-dd-redflags","page":"/legal/ma-dd-redflags","json":"/use-cases/ma-dd-redflags.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"tariff-classification","num":"70","name":"Tariff classification memo","status":"live","industries":["compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Memo top-1 subheading (6-digit)","value":"145 / 200","unit":null,"n":200,"split":"test","note":"held-out CBP rulings, run once; the index excludes them and their near duplicates"},{"name":"Memo top-1 heading (4-digit)","value":"171 / 200","unit":null,"n":200,"split":"test","note":null},{"name":"Memo top-3 subheading","value":"172 / 200","unit":null,"n":200,"split":"test","note":"proposal plus up to two alternatives"},{"name":"Memo top-3 heading","value":"187 / 200","unit":null,"n":200,"split":"test","note":null},{"name":"Proposed (not sent to a broker)","value":"108 / 200","unit":null,"n":200,"split":"test","note":"subheading right on 81% of these"},{"name":"Same model without retrieval, top-1 subheading","value":"57 / 200","unit":null,"n":200,"split":"test","note":"baseline: description and the GRI only"},{"name":"Rulings' vote, hybrid + rerank, top-1 subheading","value":"118 / 200","unit":null,"n":200,"split":"test","note":"baseline: no model, the codes of the top 5 rulings"},{"name":"Rulings' vote, BM25, top-1 subheading","value":"116 / 200","unit":null,"n":200,"split":"test","note":"baseline: no embedder, no reranker"},{"name":"Ruling quotes verified word for word","value":"415 / 429","unit":null,"n":429,"split":"test","note":"unverified quotes send the memo to a broker"}],"dataset":"CBP CROSS New York classification rulings since 2018 (public domain), 35 headings in 11 chapters: 200 held-out test and 60 dev rulings; the index holds the other rulings minus near duplicates.","held_out":true,"caveats":["Queries are CBP's own product descriptions from the rulings, cut before the classification paragraphs and with codes masked; real product sheets are vaguer.","Gold is the code CBP gave, which can be older than the 2026 HTS release; accuracy is scored at 4 and 6 digits only.","One subset of CROSS (35 headings); rulings on the same product line by the same requester can remain in the index when their descriptions differ.","A broker's judgment on the memos that went to a broker was not measured."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/tariff-classification"},"quality_evidence":[{"tier":"lite","label":"Lite · smaller reranker","evidence":[{"metric":"held-out accuracy with the 0.6B reranker","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo","evidence":[{"metric":"memo top-1 subheading (6-digit), held-out rulings (n=200, run once)","value":"145 / 200","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"},{"metric":"memo top-1 heading (4-digit), held-out","value":"171 / 200","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"},{"metric":"memo top-3 subheading, held-out","value":"172 / 200","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"},{"metric":"memos proposed (not sent to a broker) and their subheading accuracy, held-out","value":"108 / 200 proposed; 81% right","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"},{"metric":"same model without retrieval, top-1 subheading, held-out","value":"57 / 200","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"},{"metric":"ruling vote alone (hybrid + rerank) / BM25 alone, top-1 subheading, held-out","value":"118 / 200 / 116 / 200","source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27"}]}],"benchmark":{"title":"How often is the proposed code right?","intro":"Held-out CBP New York rulings: the product description from each ruling (tariff numbers masked, the classification paragraphs cut) against the code CBP gave. The index never holds the held-out rulings or near duplicates of them. Prompts and checks were set on 60 dev rulings; the 200 test rulings were run once.","rows":[{"label":"Memo, top-1 heading / subheading","value":"86% / 72%","detail":"n = 200; top-3 subheading 86%"},{"label":"Proposed memos only","value":"81% subheading right","detail":"108 of 200 proposed; the rest sent to a broker with the reason"},{"label":"Same model, no retrieval","value":"50% / 28%","detail":"heading / subheading, top-1"},{"label":"Rulings' vote, no model (hybrid + rerank / BM25)","value":"59% / 58%","detail":"subheading, top-1"}],"points":[{"heading":"Retrieval is most of the gain","text":"The same model with no rulings to read gets the subheading right far less often; with them it beats the rulings' own vote because it reads the product against the heading texts."},{"heading":"It declines a lot","text":"Many memos go to a broker, usually because the description lacks a fact the line turns on (fibre shares, gender, essential character) or a quote could not be verified. That is the intended direction of error, and it costs coverage."}],"source":"decosa-api docs/evals/tariff-classification.md, 2026-09-27; docs/evals/tariff-classification/test-score.json"},"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":32130,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.0041},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone into a clean directory, the api image built from it, compose up (named volume), rehearsal bundle and the heater sample against the already-running local Qwen3.8-27B (direct route) and decosa-retrieval services","notes":"9/9 rehearsal checks with the bundled 198-ruling sample set; receipts attested. The retrieval container build and the one-hour ruling fetch in the assemble prompt were not re-run in the sandbox (no new GPU loads; the fetch ran on the host)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.","The eval asks about products CBP already ruled on, described in CBP's own words; real product sheets are vaguer, so expect more 'needs a broker'.","Older rulings cite statistical numbers that no longer exist; the memo keeps the valid 8- or 6-digit prefix and flags it.","1,992 rulings in 35 headings; a product outside those chapters gets thin evidence and goes to a broker."],"receipt_coverage":"full"},"cost_per_run_usd":0.0041,"rehearsal_bundle":{"url":"/samples/tariff-classification.zip","checks":9,"bytes":2320},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"One call per memo: reads the rulings found, the HTS text of the candidate headings and the GRI, and proposes a heading, subheading and statistical number with quotes, or declines.","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Qwen3-Embedding-0.6B","role":"Embeds the ruling set once (cached) and each product description, for the dense half of the hybrid search (BM25 is the other half).","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Embedding-0.6B"},{"name":"Qwen3-Reranker-4B","role":"Scores the 40 best chunks for each description, so the rulings the model reads are the closest products, not the closest words.","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Reranker-4B"},{"name":"Checks and record (decosa-api, Python)","role":"Codes against the HTS release, codes nest, quotes word for word with byte offsets, a cited ruling at the proposed heading, rulings agree, evidence not thin; the signed record.","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/tariff-classification","page":"/tools/finance/tariff-classification","json":"/use-cases/tariff-classification.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"csr-number-verifier","num":"71","name":"CSR number-to-table verifier","status":"live","industries":["science-research","healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted number errors caught","value":"92 / 96","unit":null,"n":96,"split":"test","note":"held-out synthetic CSRs, 8 errors each: transposition, other arm, wrong N, wrong denominator, rounding, stale cut; re-measured 30 Sep 2026"},{"name":"Planted number errors caught, second writer","value":"93 / 96","unit":null,"n":96,"split":"test","note":"the held-out CSRs with each paragraph rewritten by Qwen3.8 (numbers kept)"},{"name":"Planted number errors caught, ClinicalTrials.gov tables","value":"170 / 178","unit":null,"n":178,"split":"heldout","note":"30 trials' posted results as CSR tables, template narrative"},{"name":"False flags per 100 numbers on clean reports","value":"0 (0 / 899)","unit":null,"n":899,"split":"test","note":null},{"name":"Traced numbers citing the gold cell","value":"99.1%","unit":null,"n":884,"split":"test","note":null},{"name":"Time per 100 pages","value":"129 s (JSON); 300 s (PDF); 493 s (scan)","unit":null,"n":null,"split":"test","note":null}],"dataset":"Synthetic CSRs of fictional phase 3 trials (sections 10-12, 11 tables each, an earlier data cut): dev seeds 1-4, held-out seeds 101-112 with a template narrative and again rewritten by Qwen3.8; 30 ClinicalTrials.gov trials' posted results laid out as CSR tables with a template narrative. Each report scored clean and with 8 planted errors.","held_out":true,"caveats":["Same author wrote the synthetic reports, the planter and the prompts; real CSRs are longer and messier.","The ClinicalTrials.gov narratives come from templates, not from a sponsor's CSR.","No real CSR and no comparison with a QC reviewer's findings.","Misses are mostly the other arm's value when nothing else in the sentence breaks (3 of 24 held out). Numbers inside real arm names (for example 'Chondroitin 4&6') cause most false flags on the ClinicalTrials.gov set."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/csr-number-verifier"},"quality_evidence":[{"tier":"lite","label":"Lite · text and rows, keyword search","evidence":[{"metric":"Planted number errors caught, held-out synthetic CSRs, keyword search","value":"93 / 96 (96.9%), 1 false flag in 899 clean numbers","source":"decosa-api docs/evals/csr-number-verifier.md, test split, BM25 ablation"}]},{"tier":"standard","label":"Standard · retrieval + document reader + the model (hosted demo)","evidence":[{"metric":"Planted number errors caught, held-out synthetic CSRs (template narrative)","value":"92 / 96 (95.8%)","source":"decosa-api docs/evals/csr-number-verifier.md, test split"},{"metric":"Planted number errors caught, held-out synthetic CSRs rewritten by a second writer (Qwen paraphrase, numbers kept)","value":"93 / 96 (96.9%)","source":"decosa-api docs/evals/csr-number-verifier.md, test-para split"},{"metric":"Planted number errors caught, 30 ClinicalTrials.gov trials laid out as CSR tables","value":"170 / 178 (95.5%)","source":"decosa-api docs/evals/csr-number-verifier.md, ctgov split"},{"metric":"False flags on clean reports (per 100 numbers checked)","value":"0 per 100 (0 / 899) on held-out synthetic; 0.11 (1 / 899) rewritten; 2.04 (19 / 931) on ClinicalTrials.gov tables","source":"decosa-api docs/evals/csr-number-verifier.md"},{"metric":"Traced numbers citing the right cell","value":"99.1% (884 numbers, held-out synthetic); 98.2% (901, ClinicalTrials.gov)","source":"decosa-api docs/evals/csr-number-verifier.md"},{"metric":"Time per 100 pages","value":"129 s from JSON; 300 s from a born-digital PDF, 493 s from a scan (reading included), shared gateway","source":"decosa-api docs/evals/csr-number-verifier.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":6525,"p95_ms":null,"runs":null,"receipts_per_run":6,"cost_per_run_usd":0.004468},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B, the running document reader (:8497) and retrieval (:8499) services, local signing; torn down after","notes":"The rehearsal bundle passed 15/15 in 62 s; every receipt attested. The reader and retrieval services were the running ones, not built from the compose file here."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.","Measured on synthetic CSRs written by the same author as the prompts and on ClinicalTrials.gov results with a template narrative; not on a real sponsor CSR or against a QC reviewer's findings.","Figures, listings and patient narratives are not read; RTF and SAS outputs must be exported to PDF or rows first.","A PDF run reads up to 40 pages; split a longer report by section.","An unflagged number is not proven right: 'matched by value only' means the value was found in one cell, not that the sentence was matched to it. A run where model calls failed is marked incomplete."],"receipt_coverage":"full"},"cost_per_run_usd":0.004468,"rehearsal_bundle":{"url":"/samples/csr-number-verifier.zip","checks":15,"bytes":14976},"models":[{"name":"decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers","role":"Finds the numbers, keeps the tables as typed cells, compares and recomputes in code, names the likely slip, checks in-text tables against their TLF, writes the QC report and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"One call per paragraph: which cell each number claims to report (by arm, row and timepoint), or which cells a derived number is computed from; a second look for numbers left unplaced; a second reading (values hidden) between two neighbouring cells","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","role":"Finds the tables a paragraph most likely reports in a long report (after the tables it names), with a signed receipt per search","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Reranker-4B"},{"name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","role":"PDFs and scans: finds and orders the regions of each page (Docling Heron layout) and reads tables as cells with spans (PaddleOCR-VL-1.6); born-digital text comes from the PDF's text layer","license":"Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"}],"licence":"permissive","links":{"metrics":"/metrics/csr-number-verifier","page":"/tools/life-sciences/csr-number-verifier","json":"/use-cases/csr-number-verifier.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"gpsr-listing-pack","num":"72","name":"GPSR listing pack","status":"live","industries":["compliance-trust","sales-marketing"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Missing Article 19 elements found, frozen","value":"15 / 19","unit":null,"n":19,"split":"test","note":"0 false. The 4 misses were products with a model number but no batch or GTIN; the pack now says 'check' there (Art. 9(5) accepts a type)."},{"name":"Missing elements found after that fix","value":"18 / 18","unit":null,"n":18,"split":"test","note":"Same test products, so no longer held out; 2 of 24 hit a gateway outage."},{"name":"Label-vs-sheet mismatches found","value":"5 / 5","unit":null,"n":5,"split":"test","note":"0 false flags on the other 19 products."},{"name":"Extracted fields right","value":"87 / 88","unit":null,"n":88,"split":"test","note":null},{"name":"Translations with no high flag (5 languages x 23 products)","value":"115 / 115","unit":null,"n":115,"split":"test","note":null},{"name":"Planted translation errors caught (6 languages)","value":"1,314 / 1,328 (98.9%)","unit":null,"n":1328,"split":"test","note":"The language pack's held-out drift set; 1,320 after later fixes."},{"name":"Planted meaning errors caught by the meaning check (negation, dropped clause, roles, entity, hedge; 6 languages)","value":"556 / 562 (98.9%); 522 as error","unit":null,"n":562,"split":"test","note":"Language pack meaning check, synthetic sentences. It also flags 52 of 192 correct translations for a look (4 as error); grammar is not checked."},{"name":"Planted translation errors caught, 5 added languages (fi, hu, el, bg, ro)","value":"1,064 / 1,078 (98.7%)","unit":null,"n":1078,"split":"test","note":"The language pack's drift set through the Qwen3.8-27B route, checker frozen; 1,068 after post-test fixes. 20 of 300 clean translations flagged when frozen (1 real), 9 after."},{"name":"Added EU languages at or above the language pack's FLORES bar","value":"13 / 15","unit":null,"n":15,"split":"test","note":"FLORES-200 devtest (1,012 sentences per language). COMET-22 89.1-91.7 on Qwen3.8-27B; Irish (73.4) and Maltese (69.0, and COMET cannot judge Maltese) are served as drafts that need a reviewer."}],"dataset":"32 synthetic products (8 dev, 24 test): a supplier sheet and a label each, with planted problems (a missing manufacturer e-mail or address, no EU responsible person for a non-EU maker, no model or identifier, a label number or model that differs from the sheet, no warnings).","held_out":true,"caveats":["The products, sheets, labels and planted problems are ours; no real supplier documents.","One fix was made after the test run (type-only identifiers); the after-fix numbers are on the same products.","The translations are checked in code for facts, and by a model for meaning (flipped negations, dropped clauses, swapped roles, wrong names); grammar is not checked and no native speaker has reviewed them.","Label photos were tested on one rendered label, not on phone photos of real labels.","The fifteen languages added on 27 Sep were checked on FLORES and on planted number errors in five of them, not on GPSR products; no native speaker has read them."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/gpsr-listing-pack"},"quality_evidence":[{"tier":"standard","label":"Standard · Qwen3.8-27B plus Hy-MT2-7B (hosted demo)","evidence":[{"metric":"Missing Article 19 elements found (24 held-out synthetic products)","value":"15 of 19 found, 0 false (frozen); 18 of 18 after one fix (22 products)","source":"docs/evals/gpsr-listing-pack.md, test split, 27 Sep 2026"},{"metric":"Label-vs-sheet mismatches found","value":"5 / 5, 0 false flags","source":"docs/evals/gpsr-listing-pack.md, test split"},{"metric":"Extracted fields right (model, GTIN, manufacturer name and e-mail)","value":"87 / 88","source":"docs/evals/gpsr-listing-pack.md, test split"},{"metric":"Planted translation errors caught (numbers, units, negations, codes, dates; 6 languages)","value":"1,314 / 1,328 (98.9%) frozen; 1,320 / 1,328 after fixes","source":"docs/evals/language-pack.md, drift test split"}]},{"tier":"best","label":"Best · the 30B translation model (not measured)","evidence":[{"metric":"FLORES-200 XCOMET-XXL (vendor)","value":"not measured yet","source":"Hy-MT2 report, arXiv 2605.22064 (vendor figure 89.83)"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":11961,"p95_ms":18088,"runs":8,"receipts_per_run":37,"cost_per_run_usd":0.004069},"selfhost":null,"known_limits":["Synthetic products only in the eval; no real supplier sheets or label photos were measured.","Harmonisation legislation (toys, electrical, cosmetics) adds label rules this pack does not check.","Translations are checked for numbers, units, dates, codes, negations and locked terms, not for wording; a native speaker should read them.","Scanned PDF sheets need the document reader block's service, which is not in the hosted demo; label photos are read by Qwen3.8."],"receipt_coverage":"partial"},"cost_per_run_usd":0.004069,"rehearsal_bundle":{"url":"/samples/gpsr-listing-pack.zip","checks":6,"bytes":2041},"models":[{"name":"decosa-api GPSR pack (decosa_api/verticals/gpsr), on the language-pack block (decosa_api/lang) and the signed record (07)","role":"Article 19 checklist, value lookup against the documents, label-vs-sheet comparison, listing export and signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4, vision tower on)","role":"Label reading (image input) and field extraction as JSON","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Hy-MT2-7B","role":"Translation of warnings and safety information (the language-pack block): German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Qwen3.8-27B (the language-pack block's route for these languages)","role":"Translation of warnings and safety information: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)","license":"Apache-2.0","hf_repo":"Qwen/Qwen3.8-27B"}],"licence":"permissive","links":{"metrics":"/metrics/gpsr-listing-pack","page":"/tools/finance/gpsr-listing-pack","json":"/use-cases/gpsr-listing-pack.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"medical-chronology","num":"73","name":"Medical chronology with page cites","status":"live","industries":["legal","healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Planted events found","value":"189 / 198","unit":null,"n":198,"split":"test","note":"6 held-out synthetic packets, 96 pages; prompts and thresholds frozen after the dev split"},{"name":"Date right, of events found","value":"189 / 189","unit":null,"n":189,"split":"test","note":null},{"name":"Cite on the right page and box, of events found","value":"189 / 189","unit":null,"n":189,"split":"test","note":"the cited box overlaps the planted line by at least half its height"},{"name":"Lines that are not planted events","value":"2 / 224","unit":null,"n":224,"split":"test","note":"both were therapy techniques in a PT plan listed as procedures"},{"name":"Planted date conflicts flagged","value":"6 / 6","unit":null,"n":6,"split":"test","note":"0 false conflicts"},{"name":"Planted gaps in treatment found","value":"6 / 6","unit":null,"n":6,"split":"test","note":"0 false gaps"},{"name":"Pre-existing conditions flagged (related or not said right)","value":"12 / 12 (12 / 12)","unit":null,"n":12,"split":"test","note":"0 false flags"},{"name":"Events on faxed copies cited on the copy too","value":"57 / 66","unit":null,"n":66,"split":"test","note":null},{"name":"Seconds per 100 pages","value":"537.7","unit":"s","n":96,"split":"test","note":"end to end on the shared hosted gateway, one packet at a time"},{"name":"Planted events found, dev","value":"98 / 99","unit":null,"n":99,"split":"dev","note":"3 packets used while building; 1 of 114 lines not a planted event"}],"dataset":"Nine synthetic record packets for fictional patients (3 dev, 6 held-out test), 16 pages each from five providers: typed PDFs with a text layer, scans with rotation, speckle and JPEG noise, handwritten forms and flow sheets, and faxed copies, with planted events, copies, one date conflict, one gap and two pre-existing conditions each (decosa_api/verticals/chronology/synth.py).","held_out":true,"caveats":["Same author wrote the generator, the answer key and the prompts; one template family, tidier than real records.","Matching an entry to a planted event uses its type and key words; a few defensible extra lines (an operative note listed as a visit, orders as referrals) are counted as neither right nor wrong.","No real records and no nurse reviewer's chronology to compare with.","Timing is on a gateway shared with other workloads; a private GPU runs faster."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/medical-chronology"},"quality_evidence":[{"tier":"lite","label":"Lite · document reader and the model, no retrieval service","evidence":[{"metric":"Planted events found, conflicts, gaps with BM25 only","value":"not measured yet","source":"the eval runner supports it (scripts/chronology_eval.py --bm25); not run on the test split"}]},{"tier":"standard","label":"Standard · reader, retrieval service and the model (hosted demo)","evidence":[{"metric":"Planted events found (held-out, 6 packets, 96 pages)","value":"189 / 198 (95.5%)","source":"decosa-api docs/evals/medical-chronology.md, test split, 27 Sep 2026"},{"metric":"Date right / cite on the right page and box, of those found","value":"dates 189 / 189, cites 189 / 189","source":"decosa-api docs/evals/medical-chronology.md, test split"},{"metric":"Lines that are not planted events","value":"2 / 224 (0.9%)","source":"decosa-api docs/evals/medical-chronology.md, test split"},{"metric":"Planted date conflicts / gaps / pre-existing flagged","value":"6 / 6, 6 / 6, 12 / 12; 0 false flags of each kind","source":"decosa-api docs/evals/medical-chronology.md, test split"},{"metric":"Events on faxed copies carrying the copy's cite","value":"57 / 66","source":"decosa-api docs/evals/medical-chronology.md, test split"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":25306,"p95_ms":null,"runs":null,"receipts_per_run":8,"cost_per_run_usd":0.004391},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running document reader and retrieval services, local signing; torn down after","notes":"The rehearsal bundle passed 12/12 in 68 s (the 16-page packet uploaded as files: conflict, gap, pre-existing, copies, cites in the index, reader receipts, record verified and failed when changed; every receipt attested) and the smoke module passed in 17.9 s. The reader and retrieval services were not built from the compose file here: the running ones were used."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route): the smoke test on one 4-page file, and the full 16-page sample from the console. The production API gets this tool when the branch merges.","Measured on synthetic packets written by the same author as the prompts, one template family; not on real records or against a nurse reviewer's chronology.","When the reader skips a line on a scan, the event on it is missed: 9 of 198 planted events were missed, mostly this way.","Hosted runs are capped at 40 pages; a page takes about 5 s on the shared gateway.","Billing records and charges are not extracted."],"receipt_coverage":"full"},"cost_per_run_usd":0.004391,"rehearsal_bundle":{"url":"/samples/medical-chronology.zip","checks":12,"bytes":965716},"models":[{"name":"decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks","role":"The event schema and page checks, copies, merging, conflicts, gaps, pre-existing, cites to chunks and bytes, the Markdown chronology and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"One call per page (the page's elements, plus the page image for a scan) listing typed events with quote and date; typed judgments for duplicate and conflict pairs and for pre-existing conditions; re-reads of regions the parser was unsure of","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Docling 2.130 with the Heron layout model (document reader block)","role":"Finds and orders the regions of every page (text, headings, tables, checkboxes) with their boxes","license":"MIT (Docling) + Apache-2.0 (weights)","hf_repo":"docling-project/docling-layout-heron"},{"name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","role":"Reads each region of a scanned page (text, handwriting, tables as cells)","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"},{"name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","role":"Scans into retrieval: each page's elements as chunks with page and box, the index and layout hashes, every cite located to a chunk and byte span, and a reranked search per procedure for other pages that give another date","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Reranker-4B"}],"licence":"permissive","links":{"metrics":"/metrics/medical-chronology","page":"/legal/medical-chronology","json":"/use-cases/medical-chronology.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"nsa-idr-packet","num":"74","name":"No Surprises Act IDR packet and eligibility screen","status":"live","industries":["healthcare","finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Eligibility verdict right from documents, held out","value":"61 of 61","unit":null,"n":61,"split":"test","note":"One model call per case reads every fact; split fixed before any run and run once; templated synthetic documents"},{"name":"Planted ineligible disputes caught on the planted check","value":"32 of 32","unit":null,"n":32,"split":"test","note":"17 kinds: late or missing negotiation, early or late initiation, state law, opt-in, consent, public coverage, ground ambulance, in network, coverage denial, cooling-off, four batching faults"},{"name":"Eligible disputes passed, including traps","value":"29 of 29","unit":null,"n":29,"split":"test","note":"Traps: ancillary consent, self-insured in state-law states, day-30 and day-34 boundaries, a moved cooling-off window, air ambulance"},{"name":"Facts read right from the documents","value":"496 of 497","unit":null,"n":497,"split":"test","note":"The miss: an issue date read as the receipt date; the verdict was still right"},{"name":"Eligibility, rules only with full facts","value":"91 of 91","unit":null,"n":91,"split":"synthetic","note":"No model; gold clocks from a separate numpy implementation"},{"name":"Planted unsupported brief sentences removed","value":"25 of 25","unit":null,"n":25,"split":"synthetic","note":"Written once before the run; 22 of 22 supported sentences kept"},{"name":"Clock comparisons agreeing with a second implementation","value":"160,000 of 160,000","unit":null,"n":160000,"split":"synthetic","note":"20,000 random dates, 8 clocks, 2022-2027"}],"dataset":"91 synthetic disputes from scripts/idr_eval_cases.py (seed 74): dev 30, test 61; three hand-written briefs with 25 planted sentences; 20,000 random dates for the clocks. All CC0.","held_out":true,"caveats":["Synthetic, templated documents with one fact per line, written by the same author as the rules and prompts; real remittances, scanned EOBs and letters are harder. These numbers do not predict accuracy on real claim files.","The rules score shows the code matches the author's reading of 45 CFR 149.510 and CMS guidance, not that the reading is right; it was not checked against CMS's recorded eligibility outcomes.","The planted-sentence set is small (47 sentences) and the planted errors are the obvious kind.","The extraction prompt was adjusted on the demo samples before the dev run; distractor dates were added after the first dev run; test was run once."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/nsa-idr-packet"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"eligibility and brief grounding","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"eligibility verdict from documents alone, held-out test (61 synthetic disputes)","value":"61/61; 32/32 planted ineligible caught on the planted check; 29/29 eligible, including traps, passed","source":"decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27, gateway route; split fixed before any run; templated synthetic documents"},{"metric":"eligibility, rules only with full facts (91 planted cases)","value":"91/91; gold clocks from a separate implementation","source":"decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27"},{"metric":"facts read right from the documents (dev / test)","value":"243/243 / 496/497 (the miss: an issue date read as the receipt date)","source":"decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27"},{"metric":"planted unsupported brief sentences removed / supported sentences kept","value":"25/25 / 22/22 (changed figures, invented credentials and statistics, contradictions, moved dates, billed charges, Medicare and UCR rates)","source":"decosa-api docs/evals/nsa-idr-packet.md, measured on our server 2026-09-27; small n"},{"metric":"business-day clocks against a second implementation (numpy over the OPM holiday list)","value":"160,000/160,000 comparisons agree","source":"decosa-api docs/evals/nsa-idr-packet.md, 2026-09-27"},{"metric":"real remittances, scanned EOBs and IDR outcomes","value":"not measured yet","source":null}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"eligibility and brief grounding","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"eligibility and brief grounding","value":"not measured yet","source":null}]}],"benchmark":{"title":"Does it catch the disputes that should not be filed?","intro":"91 synthetic disputes from a generator with a structured truth: 40 eligible (24 of them traps that look ineligible) and 51 planted ineligible, 3 of each of 17 kinds. The gold clocks come from a separate implementation. In the documents run the model reads every fact from EOB and notice text with distractor dates; a 61-case test split was fixed before any run and run once.","rows":[{"label":"Held-out test, verdict right from documents","value":"61 of 61","detail":"32 ineligible caught on the planted check, 29 eligible passed"},{"label":"Facts read right from documents","value":"496 of 497","detail":"test; the miss read an issue date as the receipt date"},{"label":"Planted unsupported brief sentences removed","value":"25 of 25","detail":"and 22 of 22 supported sentences kept"},{"label":"Cost","value":"about $0.001 per screen, $0.02 per packet","detail":"1,662 tokens (1 call) and 52,263 tokens (23 calls) on the samples, gateway list price"}],"points":[{"heading":"What it does not show","text":"The documents are templated synthetic text with one fact per line; real 835 remittances, scanned EOBs and letters are much harder to read. The same author wrote the cases and the rules, so a perfect rules score shows the code matches that reading, not that the reading is right. Check it on your own closed disputes before relying on it."},{"heading":"Where the rules are still moving","text":"The 2026 final rule's new open negotiation and initiation steps apply only after the Departments announce each IDR Gateway function; the batching changes apply to negotiations starting on or after 1 Nov 2026. Every check shows which text it applies and when it was read."}],"source":"decosa-api docs/evals/nsa-idr-packet.md, 27 Sep 2026"},"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":52100,"p95_ms":null,"runs":null,"receipts_per_run":23,"cost_per_run_usd":0.0199},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke tests passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): window 2026-10-06 to 2026-10-09, the late case failed on initiation with no model, er-anesthesia-pa likely eligible with a brief at 198.48% of the QPA in 20 s (23 attested receipts), late-initiation-az and plan-consent-defense likely ineligible with no brief, record verified; the rehearsal bundle passed 15/15. Model-server startup was not re-run."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.","Measured on 91 synthetic, templated disputes written by the building agent; not on real remittances or against CMS's recorded eligibility outcomes.","State law comes from CMS's chart (state information current as of 11 Jan 2023); a state-regulated plan in one of the 21 listed states is marked needs review unless the file says whether the state law applies.","The pre-1 Nov 2026 similar-condition batching test is left to the certified IDR entity; the CPT section ranges for the 2026 test are our approximation (the guidance was not found).","The 2026 final rule's new portal steps are not modelled yet; the clocks will need an update when the Departments announce them."],"receipt_coverage":"full"},"cost_per_run_usd":0.0199,"rehearsal_bundle":{"url":"/samples/nsa-idr-packet.zip","checks":15,"bytes":4549},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the claim facts from the EOB, remittance and notices (each with a quote), gathers the evidence for each factor the arbiter must consider, drafts the offer brief, and judges every brief sentence (the grounding judge)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/nsa-idr-packet","page":"/clinics/nsa-idr-packet","json":"/use-cases/nsa-idr-packet.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"walkthrough-to-quote","num":"75","name":"Walkthrough-to-quote","status":"live","industries":["field-trades"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted items found","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Scope work, damage, hazards and access, as a line or a condition. Dev: 16 / 16."},{"name":"Lines and conditions matching a planted item","value":"36 / 37","unit":null,"n":37,"split":"test","note":"One false condition: broken stair treads on an intact stair."},{"name":"Basis right (seen / said / seen and said)","value":"26 / 33","unit":null,"n":33,"split":"test","note":"30 / 33 if the mover's 'everything in these rooms goes' counts as saying the furniture."},{"name":"Said-only items wrongly marked seen","value":"2 / 5","unit":null,"n":5,"split":"test","note":null},{"name":"Cited time range overlaps the truth shot","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Keyframe inside the shot: 29 / 33"},{"name":"Video areas and lengths: truth inside the range","value":"2 / 5","unit":null,"n":5,"split":"test","note":"Median error of the midpoint 68.7%. Dev: 2 / 6, 13.2%."},{"name":"Counts from the video exactly right","value":"11 / 14","unit":null,"n":14,"split":"test","note":null},{"name":"Priced amounts that re-compute exactly","value":"28 / 28","unit":null,"n":28,"split":"test","note":"Integer cents; subtotals 4 / 4"},{"name":"Seconds per minute of video","value":"61.5","unit":"s","n":4,"split":"test","note":"Median, gateway shared with other workloads"}],"dataset":"6 synthetic narrated job walkthroughs (rendered rooms and a yard, 33-246 s; painting, flooring, remodel, restoration, landscaping, moving) with 49 planted items and ground-truth quantities from the scene geometry; 2 dev (the demo samples), 4 held out.","held_out":true,"caveats":["The same author wrote the scenes, the ground truth and the prompts; the two dev walkthroughs are the demo samples.","Rendered rooms are cleaner than phone footage and the narration is a clear synthetic voice; real walkthroughs are not measured.","One run per held-out clip; after the first test run the eval's room matcher was fixed and one dev-motivated change made (visible damage always gets a line); both runs are published.","A draft for the estimator; it does not measure, and its video areas are often far off."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/walkthrough-to-quote"},"quality_evidence":[{"tier":"lite","label":"Lite · scope and quote from the video only","evidence":[{"metric":"held-out test: planted items found / cited range overlaps the truth shot","value":"33 of 33 / 33 of 33 (with the narration; the video survey is the same call)","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route"}]},{"tier":"standard","label":"Standard · narration, scope and quote (hosted demo)","evidence":[{"metric":"held-out test, 4 walkthroughs: planted items found / lines and conditions that match a planted item","value":"33 of 33 / 36 of 37","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route"},{"metric":"held-out test: basis right (seen / said / seen and said) / said-only items wrongly marked seen","value":"26 of 33 / 2 of 5","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"},{"metric":"held-out test: video areas and lengths with the truth inside the range / median error; counts exact","value":"2 of 5 / 68.7%; 11 of 14","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"},{"metric":"all six clips: priced lines whose amount re-computes exactly; spoken quantities used exactly (dev)","value":"41 of 41; 3 of 3","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":54000,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0202},"selfhost":{"date":"2026-09-27","result":"pass","method":"Fresh clone of the pre-release branch into a clean directory, docker build of the api image (39 s), the api with a named volume on host networking against the local vLLM (video on) and diarizer, direct route; then torn down.","notes":"The rehearsal bundle passed 11 of 11 in 28.6 s; the flooring sample with speech to text inside the container drafted 12 lines (8 priced, $4,037.00-$4,453.50) in 38 s with 4 attested receipts and a model-call receipt for the ASR; the quote record verified; no line, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.","Measured on six synthetic walkthroughs of rendered rooms, made and labelled by the building agent; real phone footage is not measured.","Video area estimates are often far off on held-out clips (2 of 5 inside the range); the draft marks them estimated and ranks them in measure first.","Runs vary: the demo painter sample came back with 8 to 10 lines across runs; on one run the ceiling stain was a condition only (since then visible damage always gets a line)."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0202,"rehearsal_bundle":{"url":"/samples/walkthrough-to-quote.zip","checks":11,"bytes":2346108},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4), video input","role":"Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance; a long segment is split into sentences with approximate times","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)","role":"Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/walkthrough-to-quote","page":"/tools/operations/walkthrough-to-quote","json":"/use-cases/walkthrough-to-quote.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"expert-to-sop","num":"76","name":"Expert-to-SOP","status":"live","industries":["field-trades","general"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted on-screen steps found","value":"21 / 25","unit":null,"n":25,"split":"test","note":"All four misses in one recording whose times came back in 2-second slots. Dev: 14 / 14."},{"name":"Found steps in the right order","value":"100%","unit":null,"n":21,"split":"test","note":null},{"name":"Cited start within 3 s","value":"19 / 21","unit":null,"n":21,"split":"test","note":"Median error 0.7 s"},{"name":"Said-but-not-shown steps flagged needs confirmation","value":"4 / 5","unit":null,"n":5,"split":"test","note":null},{"name":"Shown-but-not-said steps marked not narrated","value":"3 / 4","unit":null,"n":4,"split":"test","note":null},{"name":"Seen steps wrongly flagged","value":"0","unit":null,"n":25,"split":"test","note":null},{"name":"Invented steps","value":"1","unit":null,"n":null,"split":"test","note":"A real click just outside the matcher's window; counted against us"}],"dataset":"6 synthetic narrated screen recordings (32-49 s) of scripted procedures in three fictional desktop apps, a stock synthetic voice, one said-only and one shown-only step planted in each; 2 dev (the demo samples), 4 held out.","held_out":true,"caveats":["The same author wrote the apps, the procedures, the ground truth and the prompts; the two dev recordings are the demo samples.","Screen recordings of simple desktop tasks with a clean synthetic voice; bench footage, long procedures and real speech are not measured.","One held-out recording's times came back in 2-second slots; a timing warning was added after the test run and changes no step or status.","A draft for a named reviewer; it does not judge whether a procedure is safe or correct."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/expert-to-sop"},"quality_evidence":[{"tier":"lite","label":"Lite · steps and keyframes, no narration check","evidence":[{"metric":"held-out test, 4 recordings: planted on-screen steps found / in order / start within 3 s","value":"21 of 25 / 100% / 19 of 21 (the draft is the same call as Standard)","source":"decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route"}]},{"tier":"standard","label":"Standard · steps checked against the narration (hosted demo)","evidence":[{"metric":"held-out test, 4 recordings: said-but-not-shown steps marked needs confirmation / shown-but-not-said marked not narrated","value":"4 of 5 / 3 of 4","source":"decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route"},{"metric":"held-out test: said-and-shown steps confirmed / seen steps wrongly flagged / invented steps","value":"16 of 21 / 0 / 1","source":"decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route"},{"metric":"held-out test: recordings where the model cited times in 2-second slots instead of reading them (a timing warning is shown)","value":"1 of 4","source":"decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route"},{"metric":"demo samples (dev, used while writing the prompts): planted steps found / flags as planted","value":"14 of 14 / 4 of 4","source":"decosa-api docs/evals/expert-to-sop.md, measured on our server 2026-09-27, gateway route"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":20200,"p95_ms":null,"runs":null,"receipts_per_run":16,"cost_per_run_usd":0.0156},"selfhost":{"date":"2026-09-27","result":"pass","method":"Fresh clone of the pre-release branch into a clean directory, docker build of the api image (32 s; ffmpeg 7.1.5 inside), the api with a named volume on host networking against the running local vLLM (Qwen3.8-27B with video input) and diarizer on the direct route; then torn down.","notes":"The rehearsal bundle passed 10 of 10 in 12.3 s; the labeldesk sample with speech to text inside the container drafted 12 steps (1 needs confirmation) in 18.6 s with 14 attested receipts and a model-call receipt for the ASR; the revision record verified; no step, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.","Measured on six synthetic screen recordings made and labelled by the building agent; bench footage and real narration are not measured.","On 1 of 4 held-out recordings the model cited times in 2-second slots instead of reading them; the timing warning catches that case, the times themselves are not corrected.","The narration check can call a paraphrase or a speech-to-text error a difference (the printer name heard as \"DocB Thermal\")."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0156,"rehearsal_bundle":{"url":"/samples/expert-to-sop.zip","checks":10,"bytes":709777},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4), video input","role":"Watches the recording (one video part, sampled at 1 frame a second) and lists each step with its time range; judges each step against the narration near its time (the grounding block's judge); lists instructions said but not drafted, then checks them back against the video","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)","role":"Clip preparation (same timeline, no audio, at most 1280 px), keyframes at each cited time, statuses, sign-off and the signed revision record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/expert-to-sop","page":"/tools/operations/expert-to-sop","json":"/use-cases/expert-to-sop.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"honest-product-imagery","num":"77","name":"Honest product imagery","status":"preview","industries":["sales-marketing","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted misrepresentations flagged","value":"48 / 48","unit":null,"n":48,"split":"test","note":"44 refused, 4 sent to a person; size, colour, text, logo, warning and feature, 8 each. Thresholds were frozen before the test ran; fixes followed (below)."},{"name":"Faithful images flagged","value":"3 / 24","unit":null,"n":24,"split":"test","note":"2 where the model boxed no product (now found by a template search), 1 carton the generator drew 12% taller. 2 / 25 after the fixes."},{"name":"Generator misrepresentations flagged (not planted)","value":"5 / 5","unit":null,"n":5,"split":"test","note":"The image model swapped a screw cap for a pump or spray top, added cartons, redrew a carton as a bottle."},{"name":"Planted flagged by Qwen3.8 alone / Claude Opus 5.5 alone","value":"46 / 47 and 47 / 47","unit":null,"n":47,"split":"test","note":"Opus 5.5 is an eval-only reference judge (same prompt and images); the colour check closed Qwen's one miss."},{"name":"Generator-garbled small text caught","value":"0 / 17 (Opus 5.5: 15 / 17)","unit":null,"n":17,"split":"synthetic","note":"Found after the test run: re-labelled by eye, prompted by Opus's flags. The main gap."},{"name":"Consent refusals","value":"7 / 7","unit":null,"n":7,"split":"synthetic","note":"No identity, bystanders, other campaign, withdrawn, expired, not enrolled, territory; each before the comparison, with a signed decision."},{"name":"Dev: planted flagged / faithful flagged","value":"24 / 24 and 0 / 12","unit":null,"n":36,"split":"dev","note":"Thresholds set here in three rounds."}],"dataset":"Six made-up products (2 dev, 4 test, split by product); 22 lifestyle scenes drawn by Wan2.2-VACE-Fun-A14B around the packs; each faithful render in three versions plus six planted misrepresentations; the raw renders hand-labelled.","held_out":false,"caveats":["Same author made the products, the plants, the checker and the labels; one image generator; synthetic products only.","Small: 48 planted and 24 faithful test images from 8 renders of 4 products, and 5 generator misrepresentations.","The test split was used after the frozen run to fix four things; those re-run numbers are not held out.","The first hand labels missed generator-garbled warnings on 5 of 13 faithful renders; the checker passed all of them.","Any person, including generator-drawn bystanders and hands, needs a consent-ledger identity: strict by design."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/honest-product-imagery"},"quality_evidence":[{"tier":"lite","label":"Lite · pixels and the model, no document reader","evidence":[{"metric":"not measured as a separate tier (the standard tier was measured)","value":"not measured","source":"decosa-api docs/evals/honest-product-imagery.md"}]},{"tier":"standard","label":"Standard · pixels, pack text and the model (hosted demo)","evidence":[{"metric":"held-out test (thresholds frozen): planted misrepresentations flagged / faithful images flagged","value":"48 of 48 (44 refused, 4 for review) / 3 of 24","source":"decosa-api docs/evals/honest-product-imagery.md, measured on our server 2026-09-27, gateway route"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":12900,"p95_ms":null,"runs":null,"receipts_per_run":10,"cost_per_run_usd":0.0018},"selfhost":{"date":"2026-09-27","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image (45 s), the api with named volumes on host networking against the running local Qwen3.8-27B (vision, direct route) and document reader, the C2PA dev certificate from the prompt's step; then torn down.","notes":"The rehearsal bundle passed 11 of 11 in 17 s (750 ml refused, a person without consent refused after two calls, a faithful image approved with IPTC metadata and a C2PA credential, the record verifies and a tampered copy fails); receipts attested; no product text in the logs. The first try without the C2PA step approved the image with no credential, which the bundle caught. The model and document-reader servers' own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route); production gets this tool when the branch merges.","Measured on six made-up products with scenes drawn by one image model; real product photos and other generators were not tested.","The alignment handles upright shots and small tilts; strong perspective or a product held at an angle skips the colour and logo checks.","A generated hand or body part counts as a person, so it needs an identity (a synthetic performer can be enrolled as a fictional identity).","The C2PA credential uses a development certificate; public validators show it as untrusted."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0018,"rehearsal_bundle":{"url":"/samples/honest-product-imagery.zip","checks":11,"bytes":413230},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Finds the product units, the logo and label, and every person or human likeness in the reference photo and the AI image; compares the two side by side with the seller's facts, per category (size, colour, text, logo, warning, feature)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Docling 2.130 with the Heron layout model","role":"Pack text, step 1: finds the text regions on crops of the product (the document reader block)","license":"MIT (Docling) + Apache-2.0 (weights)","hf_repo":"docling-project/docling-layout-heron"},{"name":"PaddleOCR-VL-1.6 (0.9B)","role":"Pack text, step 2: reads each region (name, variant, claims, quantity, warnings)","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"},{"name":"decosa-api imagery module (decosa_api/verticals/imagery)","role":"Alignment (NCC on edges over scale and a few degrees of rotation), colour (CIEDE2000 after white balance on the pack's neutral areas), logo match, unit count, changed pack area, the quantity and text comparisons, the verdict rules, the consent gate, XMP, the burned-in label and its OCR read-back, the C2PA credential and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/honest-product-imagery","page":"/tools/media/honest-product-imagery","json":"/use-cases/honest-product-imagery.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"pv-intake","num":"78","name":"Pharmacovigilance intake","status":"live","industries":["healthcare","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Missing minimum criteria caught","value":"17 / 20","unit":null,"n":20,"split":"heldout","note":"Run once. Reporter 2 of 5 (anonymous consumers passed); patient, product and event 5 of 5 each. A rule change found on this set makes it 20 of 20 (not held out)."},{"name":"Criteria wrongly called missing","value":"1 / 300","unit":null,"n":300,"split":"heldout","note":"'The mother of a teenage girl' not read as identifying the patient."},{"name":"Seriousness right","value":"60 / 60","unit":null,"n":60,"split":"heldout","note":"Valid reports with an event; case level. Criteria set exact 59 of 60."},{"name":"Expectedness right","value":"56 / 60","unit":null,"n":60,"split":"heldout","note":"3 false unexpected, 1 unclear. Blind frontier judge (Claude Code Opus sub-agent): 59 of 60."},{"name":"Day 0 exact","value":"55 / 59","unit":null,"n":59,"split":"heldout","note":"Misses: a vendor-sent report (2), a date without a year, a day-first date."},{"name":"Deadlines exact, end to end","value":"78 / 86","unit":null,"n":86,"split":"heldout","note":"US 15-day, EU 15- and 90-day; misses follow day 0 and expectedness errors."},{"name":"Clock code against the labeller's dates","value":"180 / 180","unit":null,"n":180,"split":"heldout","note":"And 60,000 of 60,000 random day-0s against a second implementation."},{"name":"Kept fields outside the gold","value":"12 / 410 (2.9%)","unit":null,"n":410,"split":"heldout","note":"5 of 410 not written in the report at all (inferred consumer, sex from a name); 72 fields dropped by the quote check."},{"name":"Reports not in English: expectedness right","value":"18 / 20","unit":null,"n":20,"split":"heldout","note":"24 reports in 7 languages; seriousness 20 of 20; validity 24 of 24."},{"name":"Voiced calls: validity right","value":"10 / 10","unit":null,"n":10,"split":"heldout","note":"MOSS-Transcribe-Diarize on house-voice calls; seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7."}],"dataset":"80 synthetic adverse event reports (24 not in English) about four fictional products, with planted missing criteria, seriousness and expectedness nuances and earlier company-awareness dates, written and labelled by a separate agent that never saw the prompts; 10 call transcripts voiced with house voices.","held_out":true,"caveats":["Synthetic reports from one labeller, not real case files or a safety physician's labels.","Run once; 80 reports, so the intervals are wide.","Two rule changes were made after the run from errors it found (anonymous consumers, day-first dates); the headline numbers are before them.","The narrative was off; calls not in English and scans beyond the demo form were not measured.","The frontier judge is a Claude Code sub-agent, blind to labels, not the API."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/pv-intake"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48-80 GB card","evidence":[{"metric":"the planted set","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo","evidence":[{"metric":"planted set, 80 reports (24 not in English), run once: missing minimum criteria caught","value":"17 of 20 run once (reporter 2 of 5; patient, product, event 5 of 5); 20 of 20 after a rule change found on this set","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"invented fields (patient and reporter fields kept that the report does not state)","value":"12 of 410 kept fields outside the gold (2.9%); 5 of 410 not written in the report at all","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"seriousness right / expectedness right (valid cases with an event, case level)","value":"seriousness 60 of 60; expectedness 56 of 60","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"day 0 exact / deadlines exact (end to end)","value":"day 0 55 of 59; deadlines 78 of 86","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"clock code against the labeller's own dates and a second implementation","value":"180 of 180 labeller dates; 60,000 of 60,000 random day-0s against a second implementation","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"not in English: seriousness / expectedness / invented fields","value":"24 reports in 7 languages: seriousness 20 of 20, expectedness 18 of 20, 2 of 132 fields outside the gold","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"calls (voiced from 10 call transcripts, transcribed by MOSS-Transcribe-Diarize)","value":"validity 10 of 10, seriousness 7 of 7, expectedness 6 of 7, day 0 6 of 7","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"},{"metric":"frontier comparison (blind Claude Code Opus sub-agent, same cases): seriousness / expectedness","value":"open 60/60 and 56/60; blind Opus 60/60 and 59/60","source":"decosa-api docs/evals/pv-intake.md, measured on our server 2026-09-27, gateway route"}]},{"tier":"best","label":"Best · a larger judge","evidence":[{"metric":"the planted set","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted: the best setup · a large judge with room to spare","evidence":[{"metric":"the planted set","value":"not measured yet","source":null}]}],"benchmark":{"title":"How well does it draft a case?","intro":"80 synthetic reports (24 not in English) run once through the hosted pipeline, labelled by a separate agent that never saw the prompts; the same seriousness and expectedness questions given blind to a Claude Code Opus sub-agent.","rows":[{"label":"Missing minimum criteria caught","value":"17 of 20","detail":"20 of 20 after a rule change found on this set"},{"label":"Seriousness / expectedness right","value":"60/60 · 56/60","detail":"blind Opus: 60/60 · 59/60"},{"label":"Day 0 exact","value":"55 of 59","detail":"clock arithmetic 180 of 180 against the labeller"},{"label":"Fields kept that the report does not state","value":"12 of 410","detail":"5 of 410 not written at all"}],"points":[{"heading":"Where it fails","text":"Anonymous consumers were first counted as identifiable reporters (fixed in the rules). Day 0 is missed when the medical information vendor sends the report, or the earlier date has no year. Expectedness slips when the model splits one event into parts or meets a near-synonym."},{"heading":"What it does not show","text":"Real case files are messier and a safety physician's labels may differ. Measure it on your own closed cases before relying on it."}],"source":"decosa-api docs/evals/pv-intake.md, 27 Sep 2026"},"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":7082,"p95_ms":92026,"runs":6,"receipts_per_run":4,"cost_per_run_usd":0.002944},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile (ffmpeg present), compose api service with a named data volume, direct route, local signing; torn down after","notes":"The assembly prompt's smoke test passed against the already-running local servers (network_mode host instead of the compose llm, mt, diarize and docreader services): clock 2026-09-17 for us_15 and eu_15; the email valid, serious, unexpected, day 0 2026-09-02, due 2026-09-17; the call transcribed by the diarizer, valid, serious, unexpected, day 0 2026-09-15, due 2026-09-30; the German email not valid (missing patient) with the follow-up question in German; the record verified (19 entries); all receipts attested. The rehearsal bundle passed 18/18. Model-server startup was not re-run."},"known_limits":["Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (29.8, 39.8, 92.0 s) and 3 on a quiet gateway on 28 Sep (7.0, 7.0, 7.1 s). Cost is the median at list price, model calls included (range $0.0028 to $0.0030).","Run once, 3 of 20 anonymous-reporter reports were passed as valid (fixed by a rule change found on the same set, so that fix is not held out).","Day 0 is missed when a medical information vendor sends the report ('our agent took the call on ...') or the earlier date has no year.","Measured on 80 synthetic reports from one labeller, not real case files or a safety physician's labels; calls not in English and Qwen3-ASR not measured; scans measured on the demo form only.","Speech recognition and the document reader share a busy GPU on the hosted demo; a call can fail to transcribe when that GPU is full (it retries, then says so)."],"receipt_coverage":"full"},"cost_per_run_usd":0.002944,"rehearsal_bundle":{"url":"/samples/pv-intake.zip","checks":18,"bytes":1099952},"models":[{"name":"decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)","role":"Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the fields with quotes, judges the seriousness criteria per event and expectedness against the label, drafts the narrative and judges its sentences (grounding)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Qwen3-ASR-1.7B (language pack)","role":"Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-ASR-1.7B"},{"name":"Hy-MT2-7B (language pack)","role":"Reports not in English, translated line by line to English, and follow-up questions back into the reporter's language","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Docling 2.130 with the Heron layout model (document reader block)","role":"Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positions","license":"MIT (Docling) + Apache-2.0 (weights)","hf_repo":"docling-project/docling-layout-heron"},{"name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","role":"Reads each region of a scanned form","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"}],"licence":"permissive","links":{"metrics":"/metrics/pv-intake","page":"/tools/life-sciences/pv-intake","json":"/use-cases/pv-intake.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"label-consistency-check","num":"80","name":"Label consistency across PI, SmPC, CCDS and carton","status":"live","industries":["healthcare","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted drifts caught","value":"31 / 31","unit":null,"n":31,"split":"test","note":"held-out invented medicine: 8 numbers, 5 missing warnings, 4 missing contraindications, 5 storage, 9 translations (German, French); all triaged likely error"},{"name":"Quotes tight to the drift","value":"30 / 31","unit":null,"n":31,"split":"test","note":"the overlapping quote is a phrase or sentence (at most 250 characters)"},{"name":"Logged twins triaged deliberate","value":"6 / 6","unit":null,"n":6,"split":"test","note":"the same drift with a deviation-log line that explains it"},{"name":"False flags on the clean set","value":"3 / 24","unit":null,"n":24,"split":"test","note":"all three from the translation number lock (German compound noun, decimal comma); 0 after a post-test fix, re-scored on the same model outputs"},{"name":"Triage agreement, Qwen3.8","value":"35 / 35","unit":null,"n":35,"split":"test","note":"deliberate or error on the test cases; 37 / 37 on dev"},{"name":"Triage agreement, blind Opus 5.5 (same prompt)","value":"35 / 35","unit":null,"n":35,"split":"test","note":"Claude Code sub-agent, inputs only; 36 / 37 on dev"},{"name":"Planted drifts caught, dev","value":"31 / 31","unit":null,"n":31,"split":"dev","note":"the medicine the prompts were tuned on (three rounds)"}],"dataset":"Two invented medicines, each a seven-document set (CCDS, US PI, SmPC in English, German and French, leaflet, carton): Norvexa for dev, Pelmora held out and run once on the frozen pipeline. 31 drifts planted one at a time per medicine, plus 6 twins with an explaining deviation-log line.","held_out":true,"caveats":["Same author wrote the documents, the plants, the prompts and the key; the drift types and conventions are the ones the prompt names.","Synthetic only: no real label set and no comparison with a labelling reviewer's findings.","Small n: 31 drifts and 6 twins in the test split; 31 / 31 still leaves a 95% lower bound near 89%.","The three test false flags were fixed after the test run; the held-out figure stays 3 / 24.","The triage comparison shows both models follow written conventions; it says nothing about unlogged differences a reviewer knows were approved."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/label-consistency-check"},"quality_evidence":[{"tier":"lite","label":"Lite · text documents, meaning check off","evidence":[{"metric":"Planted drifts caught","value":"not measured as a separate tier","source":"the eval ran with the meaning check on; 2 of 9 translation drifts on the test split were found only by the meaning check"}]},{"tier":"standard","label":"Standard · Qwen3.8-27B, Hy-MT2-7B and the document reader (hosted demo)","evidence":[{"metric":"Planted drifts caught, held-out synthetic set (numbers, missing warnings and contraindications, storage, translations)","value":"31 / 31, all triaged likely error; 30 / 31 quoted to a phrase or sentence","source":"decosa-api docs/evals/label-consistency-check.md, test split (Pelmora, run once)"},{"metric":"Same drift with an explaining deviation-log entry, triaged likely deliberate","value":"6 / 6","source":"decosa-api docs/evals/label-consistency-check.md, test split"},{"metric":"False flags on the clean held-out set","value":"3 of 24 flags (all from the translation number lock; 0 after a post-test fix, re-scored on the same model outputs)","source":"decosa-api docs/evals/label-consistency-check.md"},{"metric":"Deliberate-or-error triage against a blind frontier judge (Claude Opus 5.5, same prompt), 72 cases","value":"Qwen3.8 72 / 72, Opus 71 / 72; 35 / 35 each on the test cases","source":"decosa-api docs/evals/label-consistency-check.md"},{"metric":"Full seven-document set, hosted gateway route","value":"415 s under shared load, 43 Qwen3.8 calls, $0.0147 at list price","source":"decosa-api docs/evals/label-consistency-check.md"}]},{"tier":"best","label":"Best · the 30B translation model (not measured)","evidence":[{"metric":"Planted translation drifts caught","value":"not measured yet","source":"not run"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":2309,"p95_ms":24267,"runs":6,"receipts_per_run":3,"cost_per_run_usd":0.001104},"selfhost":{"date":"2026-09-27","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, language pack and document reader, local signing; torn down after","notes":"The rehearsal bundle passed 7/7 in 13.8 s; every receipt signed. The language pack and reader were the running services, not built from the compose file here."},"known_limits":["Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (18.2, 21.0, 24.3 s) and 3 on a quiet gateway on 28 Sep (2.1, 2.3, 2.3 s). Cost is the median at list price, model calls included (range $0.0011 to $0.0011).","Measured on synthetic label sets written by the same author as the prompts; not on a real company's labels or against a labelling reviewer's findings.","Full seven-document sets take minutes on the shared gateway (415 s measured); the Watch replay shows a recorded real run.","QRD headings are checked in English, German and French only; other EU languages get the number lock and meaning check.","Interactions, pregnancy sections, pharmacology and artwork layout are not compared."],"receipt_coverage":"full"},"cost_per_run_usd":0.001104,"rehearsal_bundle":{"url":"/samples/label-consistency-check.zip","checks":7,"bytes":7970},"models":[{"name":"decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)","role":"Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4, vision tower on)","role":"One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Hy-MT2-7B","role":"Back-translation of each translated paragraph into English for the meaning check (the language-pack block): German, French and the other languages it serves","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","role":"Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per line","license":"Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"}],"licence":"permissive","links":{"metrics":"/metrics/label-consistency-check","page":"/tools/life-sciences/label-consistency-check","json":"/use-cases/label-consistency-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"storefront-accessibility-pass","num":"89","name":"Storefront accessibility pass","status":"live","industries":["sales-marketing","compliance-trust"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planted issues caught","value":"122 / 127","unit":null,"n":127,"split":"test","note":"17 issue types on 30 pages of four test products; axe-core 52/52, keyboard and form passes 18/18, model 52/57. Misses: two wrong alts, three links named \"Details\"."},{"name":"Clean items flagged","value":"0 / 207","unit":null,"n":207,"split":"test","note":"Good alts, clear names and labelled fields on the same pages; advisories do not count."},{"name":"Clean pages with any finding","value":"0 / 7","unit":null,"n":7,"split":"test","note":null},{"name":"Alt judgement: bad alts caught / good alts flagged","value":"111 / 113 and 0 / 38","unit":null,"n":151,"split":"test","note":"Wrong product, colour swapped, file name, vague, keyword-stuffed or empty-but-needed; colour swaps 16/18."},{"name":"Blind comparison, alt cases right: Qwen3.8 / Opus 5.5","value":"57 / 60 and 59 / 60","unit":null,"n":60,"split":"test","note":"Claude Code Opus 5.5 as a blind sub-agent with the same instructions; Qwen's 3 misses are plain background photos it wants an alt for."},{"name":"Blind comparison, planted names caught: Qwen3.8 / Opus 5.5","value":"10 / 12 and 11 / 12","unit":null,"n":12,"split":"test","note":"Both left all 28 clean names alone."},{"name":"Dev: planted caught / clean items flagged","value":"73 / 74 and 0 / 119","unit":null,"n":193,"split":"dev","note":"Prompts and empty-alt rules set here in three rounds."}],"dataset":"Harbor & Pine, a made-up shop: six invented products (2 dev, 4 test, split by product) plus home, cart and checkout pages; 50 pages with 201 planted issues of 17 types and 326 clean items; alt variants per image.","held_out":true,"caveats":["Same author built the shop, the plants, the checker and the labels; one synthetic template, no real themes.","The demo scenarios use two test-split products and were run during development; the eval pages were not.","Small n for some types: keyboard trap 2, unlinked error 2, label mismatch 1.","The frontier comparison is one blind sample of 100 cases, scored by us.","Latency was measured on a shared, loaded gateway."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/storefront-accessibility-pass"},"quality_evidence":[{"tier":"lite","label":"Lite · rules, keyboard and forms, no model (CPU)","evidence":[{"metric":"held-out test: planted issues found by the rule engine and the keyboard and form passes (no model)","value":"70 of 70 of those types; the 57 model-layer issues are not checked","source":"decosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route"}]},{"tier":"standard","label":"Standard · rules, keyboard, forms and the model (hosted demo)","evidence":[{"metric":"held-out test (frozen): planted issues caught / clean items flagged / clean pages with any finding","value":"122 of 127 / 0 of 207 / 0 of 7","source":"decosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route"},{"metric":"blind comparison on 60 alt and 40 name cases: Qwen3.8-27B vs Claude Code Opus 5.5","value":"alt 57/60 vs 59/60; planted names 10/12 vs 11/12","source":"decosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":1847,"p95_ms":13303,"runs":6,"receipts_per_run":2,"cost_per_run_usd":0.000417},"selfhost":{"date":"2026-09-27","result":"pass","method":"Fresh clone of the branch into a clean directory, docker build of the api image with WITH_BROWSER=1, the api on host networking against the running local Qwen3.8-27B (vision, direct route); then torn down.","notes":"The rehearsal bundle passed 8 of 8 in 11 s; the three-page sample shop gave 13 findings (5 P1) and the clean shop none, receipts attested; a URL audit of a local http staging page (DECOSA_TESTRUNS_TARGETS + ALLOW_PRIVATE, 390 px) found its 6 issues with only GET requests reaching the server. The model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (10.7, 10.9, 13.3 s) and 3 on a quiet gateway on 28 Sep (1.8, 1.8, 1.8 s). Cost is the median at list price, model calls included (range $0.0004 to $0.0004).","Measured on one made-up shop template; real themes with lazy-loaded images, cookie banners and third-party widgets were not tested.","The model asks for an alt on a product photo used as a background behind real text; the page audit reports that as an advisory, not a finding.","Pages behind a login can only be audited self-hosted or pasted as HTML.","The demo console audits the sample shop and pasted HTML; live URLs need an API key and a verified domain."],"receipt_coverage":"partial"},"cost_per_run_usd":0.000417,"rehearsal_bundle":{"url":"/samples/storefront-accessibility-pass.zip","checks":8,"bytes":59560},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Looks at each image with its alt text (right, wrong, poor, needed but empty, decorative), reads offer text drawn into banners, judges link and button names in their context and the form's error messages","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"axe-core 4.13.0 in headless Chromium (Playwright 1.58)","role":"Opens each page in a private headless Chromium (one per audit), runs axe-core, the keyboard pass (Tab order, Enter on add-to-cart, visible focus, Escape from dialogs) and the empty-submit form pass; blocks every request that is not GET or HEAD","license":"MPL-2.0 (axe-core) + Apache-2.0 (Playwright) + BSD-3-Clause (Chromium)","hf_repo":null},{"name":"decosa-api a11y module (decosa_api/verticals/a11y)","role":"Merges the rule, keyboard, form and model results into findings ranked P1 (blocks a purchase) to P4, cites the WCAG 2.2 success criteria, writes a code fix per finding and seals the signed record","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/storefront-accessibility-pass","page":"/tools/operations/storefront-accessibility-pass","json":"/use-cases/storefront-accessibility-pass.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"fill-and-stop","num":"97","name":"Fill a form from your papers","status":"preview","industries":["general","finance"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"PDF forms: wrong answers / answers written","value":"1 / 491 (0.20%)","unit":null,"n":491,"split":"test","note":"50 held-out synthetic PDF forms (36 fillable, 8 flat, 6 scanned), gateway route; 95% CI 0.04-1.14%. The one: a previous street address written with its city. Scored from the output PDF, not the model."},{"name":"PDF forms: answers the papers give, filled","value":"490 / 497 (98.6%)","unit":null,"n":497,"split":"test","note":"6 left blank: a parent and child with the same name made one form ambiguous (5), and one answer cited the wrong line (blocked)."},{"name":"PDF forms: signature, signing-date and certification boxes left alone","value":"150 / 150","unit":null,"n":150,"split":"test","note":"Required 100%. Checked again at the write step, whatever a caller passes."},{"name":"PDF forms: blanks with the right reason","value":"409 / 409","unit":null,"n":409,"split":"test","note":"'Your papers do not say', 'you fill this' (ID, bank), 'only you can sign or confirm'. The consent rule for Spanish and French was added after the first test run (it changed reasons, not answers)."},{"name":"Blank boxes found on flat PDFs (vector / scanned)","value":"144 / 144 and 114 / 114","unit":null,"n":258,"split":"test","note":"Labels right on 144 / 144 and 113 / 114; a blank we cannot name is listed and never written. Synthetic forms from our own generator: real flat forms will be harder."},{"name":"Wrong values / values typed","value":"6 / 1,166 (0.51%)","unit":null,"n":1166,"split":"test","note":"40 fictional forms in 5 languages x 3 source bundles (120 runs); 95% CI 0.24-1.12%. Re-run 28 Sep on the shared engine, gateway route; before the move it was 5 / 1,167. 2 prior street addresses with the city, 2 dates of birth, 1 phone, 1 comment."},{"name":"Sourced values filled","value":"1,160 / 1,175 (98.7%)","unit":null,"n":1175,"split":"test","note":"10 left empty that a document gave, 3 wrong."},{"name":"Fields correctly left for the person","value":"647 / 649","unit":null,"n":649,"split":"test","note":"Fields no document answers, sensitive fields, signatures and declarations."},{"name":"Forms sent by the agent","value":"0 / 120","unit":null,"n":120,"split":"test","note":"And 120 of 120 forced submits (a click plus a script submit after the run) held by the network layer."},{"name":"Planted-injection forms: leaks / sent","value":"0 / 20 and 0 / 20","unit":null,"n":20,"split":"test","note":"18 flagged; the 2 script attacks (a beacon to another host, an auto-submit) were stopped at the network layer."}],"dataset":"PDF forms: 62 synthetic PDFs from scripts/formfill_suite_gen.py (seeded; 12 dev, 50 test): claim, benefits, W-9 style and school forms in English, Spanish and French, fillable or flat (vector or scanned). Web forms: 50 fictional forms in 5 languages x 3 synthetic source bundles plus 20 planted-injection variants (scripts/fillstop_suite_gen.py); MiniWoB++ (MIT) for the browser engine.","held_out":true,"caveats":["The same author wrote the PDF generator, the forms, the detector and the checker; real forms have messier layouts. On three real federal PDFs (IRS W-9 and 8822, GSA SF-95) most labels read well after we improved the caption rules, but some still did not (not scored).","4 of 5 wrong answers in the first PDF test run were our own label errors (a letter gave the address the truth said it did not); labels fixed and rescored, then the whole split re-run with the shipped code.","The same author wrote the generator, the forms and the checker; the forms are synthetic and short-valued (median 15 fields).","Uploads were not scored (no bundle has a document that is itself the invoice).","MiniWoB++, injection and timing runs used the direct model route (gateway out of credit); the test split used the gateway.","No human timed a manual fill; the typing-time comparison is a keystroke-level estimate.","Two blind cold-user reviews (Claude Code sub-agents playing a claimant and a security lead) read one run; no real users yet."],"date":"2026-09-28","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · text and PDF sources, English review","evidence":[{"metric":"planted form suite, test split (text and PDF sources): wrong values / values typed; submit presses","value":"6 / 1,166 (0.51%); 0","source":"decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28, gateway route"}]},{"tier":"standard","label":"Standard · adds scans, photos and a review in 24 languages (hosted demo)","evidence":[{"metric":"Harbor Mutual demo (39 fields, two PDFs and a phone photo, Spanish review): fields filled / left for the person; flag caught; submits","value":"33 / 6; yes; 0","source":"decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28"},{"metric":"20 planted-injection forms: leaks; submit presses; flagged","value":"0; 0; 18 of 20 (the other 2 were script attacks stopped at the network layer)","source":"decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28"}]},{"tier":"wanted","label":"Wanted: the best setup · a second model on choices and widgets","evidence":[{"metric":"not measured yet","value":"not measured yet","source":"not run"}]}],"benchmark":{"title":"How well does the agent use a browser on its own?","intro":"The shared computer-use engine's hybrid observer (an element table and a screenshot in one call, values only from the task text), which fill-and-stop uses for widgets, on MiniWoB++'s 84 DOM-form tasks x 5 test seeds. Qwen3.8-27B, temperature 0, through the gateway.","rows":[{"label":"Shared engine, typed values only from the task text (the product's rule)","value":"73.1% (68.7-77.1%)","detail":"307 / 420 episodes, 28 Sep 2026 re-run after fill-and-stop moved onto the engine"},{"label":"Fill-and-stop's own observer before the move (same episodes)","value":"76.2% (71.9-80.0%)","detail":"320 / 420; paired difference -3.1 points, 95% CI -6.9 to +0.2 (not significant), mostly date pickers"},{"label":"Earlier baseline, same group: element table + guards / pixels only","value":"36.7% / 61.0%","detail":"25 Sep 2026 bench"}],"points":[{"heading":"Where it fails","text":"Tasks that need a value read off the page (find a word, read a table, do the sum) are blocked by design under the product rule: page text is never a source, which is what stops an injected page from choosing what gets typed. Long searches (book a flight, the 8th search result) and a few custom widgets still fail."}],"source":"decosa-api docs/evals/fill-and-stop.md, 28 Sep 2026"},"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":13700,"p95_ms":17100,"runs":3,"receipts_per_run":3,"cost_per_run_usd":0.0025},"selfhost":{"date":"2026-09-28","result":"pass","method":"Fresh clone of the branch, docker build of the api image (default, no browser), the api on host networking against the running local Qwen3.8-27B (direct route) and document reader; then torn down (container, volume, image, clone).","notes":"PDF mode: the Harbor Mutual sample 3 times, 8.3, 8.4 and 9.9 s, 31 of 39 boxes written each time, $0.0029 at list price, the record verified by the box's own /record/verify; the copy list for the school meal sample (site class forbidden); the site table (login.gov: identity, copy). The live web-form engine was verified self-hosted with the browser build earlier on 28 Sep (rehearsal 12 of 12)."},"known_limits":["PDF answers come from one line of your papers (or the next); an answer that combines lines is left for you.","Flat and scanned PDFs: we write on the page where we found the blank; a blank we could not place or name is listed on the source sheet and left empty. Tested on synthetic flat forms; real scans vary.","Some PDF viewers do not show letters outside Western European sets in fillable boxes (the value is stored; check it on screen).","The copy list needs the website's labels (pasted, or read from a screenshot); it does not see hidden fields or later pages until you paste them.","Numbers come from branch runs on our server (28 Sep 2026). The planted and injection splits and MiniWoB++ were re-run through the gateway (receipted) after the move onto the shared engine; the federal-form timing and the last demo recording used our server's Qwen3.8-27B directly (same weights, receipts 'unverified').","A value must come from one line (or the next): an answer that combines several lines (an insurer's name, address and policy number in one box) is left for you.","Choices it infers (a claim type from a description, 'Yes' to a police report because a report exists) are marked 'check it'; the citation then points at supporting text, not a literal answer.","It does not gather facts for you: the time saved is the typing and cross-checking, not the reading. Human review time was not measured.","Blocking the commit also blocks a site's 'Save draft' and autosave while the agent works.","Hosted: demo forms and pasted HTML only; logged-in sites are the self-host command-line path (a visible browser on your machine).","The review translation is machine translation (the language pack); short labels can still be mistranslated."],"receipt_coverage":"full"},"cost_per_run_usd":0.0025,"rehearsal_bundle":{"url":"/samples/fill-and-stop.zip","checks":12,"bytes":2740},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Maps each form field to the line of your documents that answers it (12 fields per call), checks page text the patterns did not settle for instructions aimed at AI agents, and reads a widget from a screenshot when the page's code cannot set it","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","role":"Reads scanned PDFs and photos of papers (a police report photographed on a phone) into numbered lines the fields can point at","license":"Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"},{"name":"Hy-MT2-7B (language pack; Qwen3.8-27B for the EU languages Hy-MT2 does not cover)","role":"Translates the review (field labels, reasons and your cited document lines) into the person's language; the values typed into the form are never translated","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Playwright 1.58 with Chromium (headless)","role":"A throwaway headless Chromium per run (no cookies, no storage, no service workers, no downloads): reads the form, types and chooses in the page, holds every request that could send the form and blocks every other host","license":"Apache-2.0 (Playwright) + BSD-3-Clause (Chromium)","hf_repo":null},{"name":"decosa-commit-detector (XLM-RoBERTa-large fine-tune, our own model)","role":"Says whether a click the model chooses would commit something (pay, delete, send, publish, submit, security) and labels the buttons left for you; the first line, with the word list and the network hold always on","license":"Apache-2.0 (base model MIT)","hf_repo":"decosaai/decosa-commit-detector-xlmr-large"},{"name":"decosa-api fill-and-stop module (decosa_api/verticals/fillstop) on the shared computer-use engine (decosa_api.cu) and the flight recorder (tool 26)","role":"Field reader, the value checks (the value must be in the cited line, or the same date, time, phone number or amount written another way), the sensitive-field rules, the commit words in 32 languages, the injection patterns, the review, the signed approval and the release of the held Submit","license":"AGPL-3.0-or-later (the computer-use engine it runs on, decosa_api.cu, is Apache-2.0)","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/fill-and-stop","page":"/tools/operations/fill-and-stop","json":"/use-cases/fill-and-stop.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"privileged-call-notes","num":"98","name":"Privileged call notes","status":"live","industries":["legal"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Memo fact recall (blind grader)","value":"89.3%","unit":null,"n":20,"split":"test","note":"Dev: 92.2%. Grader: Claude Code Opus 5.5, blind, saw only script, ground truth and memo."},{"name":"Memo lines invented / with a wrong detail","value":"0 / 12 of 796","unit":null,"n":796,"split":"test","note":"7 of the 12 in one call (a year the call never said); the no-model year check now flags them"},{"name":"Conflicts-name recall, strict / phonetic","value":"89.6% / 94.0%","unit":null,"n":20,"split":"test","note":"Strict misses are mostly speech-recognition spellings; no invented names"},{"name":"Dated deadlines found","value":"100%","unit":null,"n":20,"split":"test","note":"Dates overall: 92.0%"},{"name":"Recording notice heard, right","value":"26/26","unit":null,"n":26,"split":"synthetic","note":"All 26 calls (dev and test). 6 calls never mention recording and are flagged"},{"name":"Time entry = call length rounded up to 0.1 h","value":"26/26","unit":null,"n":26,"split":"synthetic","note":"All 26 calls (dev and test). Narrative acceptable 20/20, client email 14/20 (blind grader)"}],"dataset":"26 synthetic attorney-client calls (6 dev, 20 test) in 9 practice areas, 3 in Spanish, voiced with Decosa house voices (Kokoro-82M) with crosstalk and a telephone band on half; run through the diarizer and the full pipeline.","held_out":true,"caveats":["Synthetic calls voiced by TTS (cleaner than real phone audio); the same model family wrote the scripts and runs the pipeline.","Fixes after cold-user tests and after the first test runs were mechanisms (spelled names, the client on the list, a weekday and year check), not tuned to test values; the year check was applied post-hoc to the final outputs.","The claim check blocks invented lines but marked only 7 of 12 lines the grader called wrong (with the year check).","Time savings are estimates from two blind persona tests, not measured with real lawyers."],"date":"2026-09-28","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card, uploaded calls","evidence":[{"metric":"call memo quality","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo","evidence":[{"metric":"memo fact recall, 20 synthetic test calls","value":"89.3%","source":"blind grader, 20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)"},{"metric":"conflicts-name recall, 20 synthetic test calls","value":"89.6% strict / 94.0% phonetic","source":"20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)"},{"metric":"invented memo items left after the claim check","value":"0 of 796","source":"blind grader on the full memo; 12 items (1.5%) had a wrong detail; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash writes and checks","evidence":[{"metric":"call memo quality","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":148000,"p95_ms":204000,"runs":26,"receipts_per_run":44,"cost_per_run_usd":0.017},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume on the host network, direct route, local signing, against the already-running Voxtral, MOSS diarizer and Qwen3.8; torn down after","notes":"Assembly prompt §6: TX/CA gives all-party; 400 without consent; custody call upload 39 s, 11 names incl. opposing counsel, 0.1 h, audio dropped, record verified (135 entries), 41 receipts attested. Live replay at 2x: 155 s, 52 captions, record verified (238 entries)."},"known_limits":["Not a certified transcript: speech recognition mishears names and numbers. Check anything that matters against the call.","The conflicts list is the names heard on the call; it does not search your conflicts database.","Deadlines are the ones said on the call. It recomputes the arithmetic, but it does not know court rules or limitation periods.","The time entry is the call audio's length rounded up to 0.1 hour; your firm's and the client's billing rules decide the entry.","The hosted demo is for synthetic calls. On the hosted API the server sees the audio and text in memory while it works; real client calls belong on your own box or the Confidential tier."],"receipt_coverage":"full"},"cost_per_run_usd":0.017,"rehearsal_bundle":{"url":"/samples/privileged-call-notes.zip","checks":10,"bytes":3199},"models":[{"name":"Voxtral Mini 4B Realtime","role":"Live captions during the call (streaming, no speakers)","license":"Apache-2.0","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602"},{"name":"MOSS-Transcribe-Diarize 0.9B","role":"After hang-up (or on an upload): who said what, one line per turn, each with its own receipt","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Speaker roles, live intake checklist, cited memo, conflicts names, claim check, time-entry narrative and client email","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/privileged-call-notes","page":"/legal/privileged-call-notes","json":"/use-cases/privileged-call-notes.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"check-their-brief","num":"99","name":"Check their brief","status":"live","industries":["legal"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Non-existent citations caught (strict)","value":"24 / 31 (77%)","unit":null,"n":31,"split":"heldout","note":"LePhantomCite eval split, run once on frozen code; 81% counting 'look at'."},{"name":"Case name and cite from two different cases (strict)","value":"34 / 68 (50%)","unit":null,"n":68,"split":"heldout","note":"75% counting 'look at'."},{"name":"Swapped-word misquotations (strict)","value":"31 / 45 (69%)","unit":null,"n":45,"split":"heldout","note":null},{"name":"Invented citations in real criticised filings called fine","value":"0 / 78","unit":null,"n":78,"split":"test","note":"30 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable (daily limit). Not held out from the fixes this run exposed."},{"name":"Misstated holdings (strict)","value":"4 / 131 (3%)","unit":null,"n":131,"split":"heldout","note":"Not reliably caught by the open judge. Our own citation-support model plus the judge: 80 / 129 at 7 / 372 false flags on held-out pairs (prototype, off by default)."},{"name":"False findings on error-free excerpts","value":"49 / 950 (5.2%)","unit":null,"n":950,"split":"heldout","note":"33 / 950 (3.5%) with the post-test fixes simulated on the same outputs."},{"name":"Hidden AI instructions flagged (blind texts)","value":"71 / 74","unit":null,"n":74,"split":"heldout","note":"The 3 misses were CJK text the planting tool could not encode. Benign hidden texts flagged: 3 / 75."},{"name":"Cost per excerpt at list price","value":"$0.0123","unit":"USD","n":390,"split":"heldout","note":"p50 9.9 s; the fictional two-page opposition: $0.0127, 32-37 s warm."},{"name":"Invented citations in criticised filings, with the local citation index (CourtListener off)","value":"32 / 78 findings, 63 / 78 findings or looks","unit":null,"n":78,"split":"test","note":"0 called fine; 4 not checked (was 32 without the index). Not held out from the lookup fixes made during that run."}],"dataset":"LePhantomCite (CC BY 4.0): 390 held-out excerpts of real federal appellate briefs with injected citation errors; 160 blind-written injection and benign texts hidden in 12 public-domain Solicitor General briefs; 31 real filings courts criticised for invented citations and 23 uncriticised briefs from RECAP.","held_out":true,"caveats":["LePhantomCite errors are injected by its authors; its non-existent citations use reporter series that do not exist, which a code check catches.","The published real-filings run had CourtListener from cache only (free-account daily limit); the re-run with the local citation index checks all but 97 of 2,030 case citations. Labels come from court orders, 17 of 31 of which list only examples.","Pin cites and misstated holdings are not reliably caught by the shipped path.","Some 'false findings' on error-free excerpts are real errors in the original briefs; not all were adjudicated.","About 5% of test items met CourtListener rate limits (shared with our other jobs).","One author wrote the rules and the dev texts; the test injection texts were blind-written.","No human paralegal was timed; the manual baseline is an estimate."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/check-their-brief"},"quality_evidence":[{"tier":"lite","label":"Lite · CPU only, no GPU","evidence":[{"metric":"Hidden-text scan and reporter/existence checks","value":"same as standard (code, no model)","source":"docs/evals/check-their-brief.md"},{"metric":"Hidden instructions without the model's reading","value":"not measured separately","source":null}]},{"tier":"standard","label":"Standard · one GPU for the judge (hosted demo)","evidence":[{"metric":"Citations to cases that do not exist, caught as a finding (strict) / as a finding or a look (lenient)","value":"24/31 (77%) / 25/31 (81%)","source":"docs/evals/check-their-brief.md, LePhantomCite held-out test (390 real brief excerpts, CC BY 4.0), run once"},{"metric":"Case name and cite that belong to two different cases","value":"34/68 (50%) strict / 51/68 (75%) lenient","source":"same"},{"metric":"Quotations with a word swapped","value":"31/45 (69%) strict / 36/45 (80%) lenient","source":"same"},{"metric":"Wrong pin cites and misstated holdings (strict)","value":"2/55 and 4/131: not reliably caught","source":"same; the judge marks most misstated holdings 'look at', as it does 25% of holdings in error-free excerpts"},{"metric":"False findings on error-free excerpts","value":"49 of 950 checked items (5.2%) as run; 33 (3.5%) with the post-test fixes simulated on the same outputs","source":"same"},{"metric":"Hidden instructions aimed at AI tools (blind-written texts hidden in 12 real briefs, 13 techniques)","value":"71/74 flagged (the 3 misses were CJK text the planting tool could not encode); benign hidden texts flagged 3/75 as run, 1/75 in the regression re-run after the pattern fix","source":"docs/evals/check-their-brief.md, section 3 (re-run after the real-PDF fixes: same numbers)"},{"metric":"Instructions written in plain view","value":"2/6 flagged, 0/5 benign flagged: weak","source":"same"},{"metric":"Real filings courts criticised for invented citations (31 filings, 161 problems from the orders; CourtListener cache-only)","value":"Invented citations: 0 of 78 called fine; 30 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable. All problems: 37/161 findings, 74/161 findings or looks","source":"docs/evals/check-their-brief.md, section 2 (Charlotin database + RECAP; not held out from the fixes it exposed)"},{"metric":"Same 54 filings with the local citation index, CourtListener off (gateway, 28 Sep 2026)","value":"Invented citations: 0 of 78 called fine; 32 findings and 63 findings or looks; 4 not checked (was 32). Case citations not checked: 97 of 2,030 (was 563). Uncriticised briefs: 47 findings on 2,006 checked items (2.3%). p50 17 s a filing.","source":"docs/evals/citation-index.md (not held out from the lookup fixes made during that run)"},{"metric":"Findings on 23 uncriticised real briefs","value":"66 of 1,645 checked items (4.0%) as run; 49 of 1,666 (2.9%) after the fixes (direct-route re-run); many are real miscites in those briefs or quotes of the other side's invalid cites","source":"docs/evals/check-their-brief.md, section 2"}]},{"tier":"wanted","label":"Wanted · a GLM-5.3-Flash holdings judge on your own hardware","evidence":[{"metric":"This eval, same protocol","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":8303,"p95_ms":8364,"runs":5,"receipts_per_run":7,"cost_per_run_usd":0.012553},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Fresh clone of a decosa-api pre-release build (6ee6b6b), the api image built from docker/api/Dockerfile (theirbrief extra), this prompt's compose with the direct route to the already-running local Qwen3.8-27B, OCR off, anonymous CourtListener. The prompt's smoke steps and the rehearsal bundle passed (11/11): 7 findings on the fictional PDF, 7 receipts, the record verifies and a tampered decision fails, the Word memo exports, a scan is refused with a clear message when OCR is off. 57.9 s, $0.0126. Torn down after."},"known_limits":["Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","No citator: it does not say whether a case is still good law.","Westlaw- and Lexis-only decisions, many unpublished orders and most state codes are not in the free sources: they come back 'look at' or 'not checkable', never 'fine'.","Case citations are looked up in a local index of CourtListener's and the Caselaw Access Project's public data (quarterly; snapshot 30 Jun 2026), with no network call. CourtListener's free API (250 searches a day) is asked only on a miss or for a volume newer than the snapshot; what nothing can answer is listed as 'not checked yet', never as fine.","Quotations and holdings of cases after about 2018 need the opinion PDF from CourtListener, one search per quoted case; if it cannot be asked, those checks stay 'not checked'.","Instructions hidden in the file are caught well; instructions written in plain view are caught only when a pattern or AI word flags the sentence first."],"receipt_coverage":"full"},"cost_per_run_usd":0.012553,"rehearsal_bundle":{"url":"/samples/check-their-brief.zip","checks":11,"bytes":4186},"models":[{"name":"decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)","role":"Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","role":"Document reader for scanned filings (no text layer): layout plus OCR, then the same checks","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"}],"licence":"permissive","links":{"metrics":"/metrics/check-their-brief","page":"/legal/check-their-brief","json":"/use-cases/check-their-brief.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"discovery-deficiency","num":"100","name":"Check their discovery responses","status":"live","industries":["legal"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Deficient responses caught on unseen real cases (RECAP test, dockets never seen in dev)","value":"174 / 235 (74%)","unit":null,"n":235,"split":"test","note":"Responses a motion to compel called deficient; 85% (174 / 205) of those the splitter found. All 43 test sets: 292 / 373 (78%)."},{"name":"Responses found by the splitter on unseen real filings","value":"75.3%","unit":null,"n":928,"split":"test","note":"27 docket-disjoint RECAP test sets; 78.8% on all 43; 99.3% on the 25 dev filings it was tuned on. Missed responses are listed as warnings."},{"name":"False flags on clean responses, blind synthetic test","value":"13 / 197 (6.6%)","unit":null,"n":197,"split":"heldout","note":"Clean responses flagged as deficient. On real filings a blind review of 80 flags the motions did not raise found 45 correct, 11 debatable, 24 wrong (15 were pre-2015 responses, since fixed)."},{"name":"Deficient-or-not precision, blind synthetic test","value":"0.897","unit":null,"n":325,"split":"heldout","note":"113 correct flags, 13 false alarms, 15 missed, 184 clean left alone; sets written by a separate blind author."},{"name":"Deficient-or-not recall, blind synthetic test","value":"0.883","unit":null,"n":325,"split":"heldout","note":null},{"name":"Per-category F1, blind synthetic test","value":"0.792","unit":null,"n":325,"split":"heldout","note":"precision 0.715, recall 0.886 before post-test fixes (0.832 after, no longer held out)."},{"name":"Recall vs blind Claude Opus 5.5, 90 responses","value":"0.952 vs 0.935","unit":null,"n":90,"split":"heldout","note":"precision 0.756 vs 0.879"},{"name":"Blind partner review: tool letter vs hand-drafted, 9 comparisons","value":"0 / 9 preferred (scores 2-5 vs 8-9)","unit":null,"n":9,"split":"test","note":"Three rounds, two of them after fixes; reviewers penalised template points instead of request-specific argument."}],"dataset":"68 sets of real written discovery responses from CourtListener RECAP (46 federal dockets; 25 dev, 43 test by hash, 27 of them on dockets with no dev set) labelled from the motions to compel; 24 synthetic sets (325 responses) written by a separate blind author; 3 synthetic demo sets (dev).","held_out":true,"caveats":["Motion labels undercount what is wrong, so precision against them is a lower bound; a blind adjudicator judged a sample of the other flags.","The adjudicator, frontier judge, letter reviewer and cold users are Claude Opus 5.5 sub-agents, not practising lawyers.","All real sets are federal; California and Texas are measured on synthetic sets only.","About a third of the real sets were rebuilt from quotes in the motion (disputed items only), and some dockets were split into correlated sets; the docket-disjoint figures are the stricter ones.","Fixes made after the held-out runs are reported separately and not counted as held out.","The drafted letter lost every blind comparison with a hand-drafted letter."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/discovery-deficiency"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card","evidence":[{"metric":"Same model and prompts as standard","value":"not measured separately","source":"estimate: identical pipeline without the document reader"}]},{"tier":"standard","label":"Standard · Qwen3.8-27B and the document reader (hosted demo)","evidence":[{"metric":"Responses correctly called deficient or not, blind synthetic test (24 sets by another author, 325 responses, federal, California, Texas)","value":"precision 0.897, recall 0.883; 184 clean responses left alone","source":"decosa-api docs/evals/discovery-deficiency.md, held out, run once"},{"metric":"Real responses a motion to compel called deficient, caught (27 docket-disjoint RECAP test sets)","value":"174 / 235 (74%); 174 / 205 (85%) of the responses the splitter found","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Clean responses wrongly flagged, blind synthetic test","value":"13 / 197 (6.6%)","source":"decosa-api docs/evals/discovery-deficiency.md, held out"},{"metric":"Splitter coverage on real filings (held out)","value":"75.3% of responses found on docket-disjoint sets (78.8% on all 43; 99.3% on the dev filings it was tuned on)","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Against a blind frontier judge (Claude Opus 5.5), 90 held-out responses","value":"recall 0.952 vs 0.935; precision 0.756 vs 0.879","source":"decosa-api docs/evals/discovery-deficiency.md"},{"metric":"Draft letter vs a hand-drafted letter, blind partner review","value":"a starting point, not ready to send: it lost 9 of 9 blind comparisons with a hand-drafted letter (2-5 vs 8-9 of 10)","source":"decosa-api docs/evals/discovery-deficiency.md"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":4756,"p95_ms":13634,"runs":5,"receipts_per_run":10,"cost_per_run_usd":0.0059},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, local signing; torn down after","notes":"The rehearsal bundle passed 11/11 in 2.8 s; the three samples ran in 2.4-3.6 s with attested receipts."},"known_limits":["The splitter finds about 3 in 4 responses in real court-filed PDFs it has not seen (75-79%); missing ones are listed as warnings.","The letter draft is a starting point, not ready to send: a blind partner review preferred hand-drafted letters on every set (9 of 9), citing missing request-specific argument.","Rule text only: no case law, local rules or standing orders. California uses the CCP 135 court calendar; Texas deadlines use the federal holiday list.","Hosted numbers are from the pre-release server before merge; production numbers follow the nightly check."],"receipt_coverage":"full"},"cost_per_run_usd":0.0059,"rehearsal_bundle":{"url":"/samples/discovery-deficiency.zip","checks":11,"bytes":3605},"models":[{"name":"decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)","role":"Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shape","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Document reader (Docling layout heron + PaddleOCR-VL-1.6)","role":"Scanned PDFs only: page images to text","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"}],"licence":"permissive","links":{"metrics":"/metrics/discovery-deficiency","page":"/legal/discovery-deficiency","json":"/use-cases/discovery-deficiency.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"payer-audit","num":"101","name":"Payer audit response","status":"live","industries":["healthcare","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Weak claims flagged, new blind BCBSM letter (29 Sep)","value":"9 / 9","unit":null,"n":9,"split":"heldout","note":"22 claims; 6 lacked an objective tool (only the client's own ratings such as SUDS). Clean flagged weak 3 / 13, all from 24-hour times written without colons; fixed after (0 / 13 on a rerun, not held out)."},{"name":"Weak claims flagged, new blind Optum letter (29 Sep)","value":"7 / 9","unit":null,"n":9,"split":"heldout","note":"A control with no measure requirement: clean flagged weak 0 / 12."},{"name":"Weak claims flagged, held-out B rerun (29 Sep)","value":"23 / 26","unit":null,"n":26,"split":"heldout","note":"Clean flagged weak 0 / 35, requirements 423 / 433; the extra miss was weak on an immediate rerun (run-to-run variance)."},{"name":"Weak claims flagged, held-out B","value":"24 / 26","unit":null,"n":26,"split":"heldout","note":"3 published policies (Evernorth BH, NC Medicaid telehealth, CMS therapy plan certification), blind-written; both misses: two-signer plans of care"},{"name":"Clean claims flagged weak, held-out B","value":"0 / 35","unit":null,"n":35,"split":"heldout","note":"pack mode"},{"name":"Requirements marked as labelled, held-out B","value":"423 / 433","unit":null,"n":433,"split":"heldout","note":"found or missing per claim per requirement"},{"name":"Respond-by date right","value":"3 / 3","unit":null,"n":3,"split":"heldout","note":"plus 5/5 on dev letters"},{"name":"Unsupported cover-letter sentences","value":"0 / 26","unit":null,"n":26,"split":"heldout","note":"blind Claude Code judge, 6 letters"},{"name":"Pasted policy instead of a pack: clean flagged weak","value":"22 / 35","unit":null,"n":35,"split":"heldout","note":"not reliable; weak 21/26"},{"name":"Weak claims flagged, dev (after fixes)","value":"38 / 39","unit":null,"n":39,"split":"dev","note":"sets A and Centene; set A was held out for v1, then used to fix mechanisms"}],"dataset":"Blind-written synthetic audits: dev = 5 sets (97 claims) on BCBSM, Centene, CMS 220.3, DME order and NC Medicaid 8C policies plus Optum; held-out B = 3 sets (61 claims) opened only after the engine was frozen.","held_out":true,"caveats":["Synthetic cases written by Claude agents; real charts are longer and messier.","v1 was run once on set A (33/34 weak but 27/48 clean flagged weak); set A then became dev, which is disclosed.","Labels are the writers' own; small n per policy.","No human auditor or practice manager rated the output yet.","Cross-claim checks (overlap, copied notes) came after the held-out design; the final engine re-run on held-out B gave the same weak and false-weak counts and no cross-claim flags.","Objective tools: BCBS Michigan does not define the term; reading it as a scored, standardised instrument's result (not a SUDS or 0-10 rating) is ours."],"date":"2026-09-29","doc_url":null},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"recommendation and criteria accuracy","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"weak claims flagged, held-out set B (pack mode)","value":"24/26; 0/35 clean claims flagged weak; both misses were plans of care with two signers (the engine reads one signature per note)","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"requirements marked as labelled, held-out set B","value":"423/433 (97.7%); the found words were on the labelled line 355/377","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"respond-by date right","value":"3/3 held-out letters (and 5/5 dev letters)","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"},{"metric":"kept cover-letter sentences a blind judge found unsupported","value":"0/26 (6 letters); 0 argued, advised or promised","source":"decosa-api docs/evals/payer-audit.md; blind judge: Claude Code (Opus 5.5) sub-agent that saw only the sources and the sentences"},{"metric":"pasted policy text instead of a pack (the model reads the requirements)","value":"weak 21/26 but 22/35 clean claims flagged weak: not reliable; review the requirements it read, or use a pack","source":"decosa-api docs/evals/payer-audit.md, held-out set B (61 synthetic claims, 3 published payer policies, cases written blind by a separate agent), engine frozen at de3d163 before the set was opened, run once on our server 2026-09-28, gateway route"}]}],"benchmark":{"title":"Does it flag the weak claims?","intro":"Blind-written synthetic audits against real published payer policies: another agent wrote each policy's requirement list, the letters, the charts with planted defects and the labels, without seeing the tool. The engine was frozen before held-out set B was opened, then run once.","rows":[{"label":"Weak claims flagged (held-out B)","value":"24 of 26","detail":"3 policies, 61 claims; both misses were plans of care with two signers"},{"label":"Clean claims flagged weak (held-out B)","value":"0 of 35","detail":"no false alarms in pack mode"},{"label":"Requirements found/missing as labelled","value":"423 of 433","detail":"held-out B, pack mode"},{"label":"Cost per audit of about 20 claims","value":"about $0.03","detail":"held-out B median $0.028, 80 s, gateway list price"}],"points":[{"heading":"Where it fails","text":"A document with two signers (a therapist's plan of care certified by a physician) is read as one signature, so a late or uncredentialed certifying signature was missed twice. Pasting the policy text instead of using a reviewed pack makes the model read the requirements, and that flagged 22 of 35 clean claims: use a pack, or review what it read."},{"heading":"What it will not do","text":"Suggest adding to, changing or back-dating a record; argue the case; judge medical necessity or coding; send anything or touch a portal."},{"heading":"What a practice manager said","text":"A blind test user playing a practice manager put a 140-chart request at about 25 hours by hand (her estimate); the tool took about 4 minutes and $0.15. She would pay $150-300 per audit once she can put her own charts in, and would still call counsel when a lot of money is at stake or the letter mentions fraud."}],"source":null},"verification":{"hosted":{"date":"2026-09-28","result":"pass","p50_ms":59900,"p95_ms":81200,"runs":5,"receipts_per_run":18,"cost_per_run_usd":0.0176},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone of the branch into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the already-running local Qwen3.8-27B, local signing; torn down after","notes":"Rehearsal bundle 10/10 in 25.3 s (4 weak claims listed first, respond by 2026-10-01, record verifies, every receipt attested); smoke ok in 21.5 s with 18/18 attested receipts and a PDF packet. Model-server startup was not re-run."},"known_limits":["Hosted timing: 5 runs of the 12-claim sample on the pre-release server over the shared gateway (47.8-81.2 s; p95 is the slowest of 5). Production is re-measured after the merge.","Synthetic cases only, written by other workloads from real published policies; no practice manager or auditor has rated the output.","One signature per note: documents signed by two people (a therapist and a certifying physician) are read as one.","Text charts only: scanned PDFs are not read yet.","A pasted policy is read by the model and is not reliable (see the eval); a reviewed pack is.","Checks across claims (overlapping sessions by the same clinician, near-identical notes) were added after the blind cold-user test; they mark claims to check by hand and are not validated on real charts (the 92% wording threshold was set on the demo data)."],"receipt_coverage":"full"},"cost_per_run_usd":0.0176,"rehearsal_bundle":{"url":"/samples/payer-audit.zip","checks":11,"bytes":6952},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the auditor's letter (who, dates, reference, policy named), reads a pasted policy's requirements when no pack is given, and for each claim points at the note's times, signature and addenda and judges each content requirement found or missing with the exact words; drafts the cover letter body and judges each of its sentences (the grounding judge)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/payer-audit","page":"/clinics/payer-audit","json":"/use-cases/payer-audit.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"prior-auth-check","num":"102","name":"Prior-auth pre-check and packet","status":"preview","industries":["healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Decision right, latest fresh set","value":"14 of 16","unit":null,"n":16,"split":"test","note":"4 real policies never used while building; run once after the cold-user fixes"},{"name":"Requests that should have waited, called ready","value":"1 of 11","unit":null,"n":11,"split":"test","note":"latest set; earlier sets 4 of 21 and 10 of 31"},{"name":"Decision right, earlier fresh set","value":"24 of 35","unit":null,"n":35,"split":"test","note":"7 real policies; run once after the guards"},{"name":"Decision right, first held-out set","value":"45 of 60","unit":null,"n":60,"split":"test","note":"12 real policies; run once before the guards"},{"name":"Undocumented read as not met","value":"0 of 111","unit":null,"n":111,"split":"test","note":"all three sets; the appeal engine's main error was 6 of 80"},{"name":"Letter sentences rated unsupported","value":"8 of 512","unit":null,"n":512,"split":"test","note":"blind reviewer, all three sets; 0 of 46 on the latest"}],"dataset":"Synthetic charts written blind against 23 real published payer policies and sections (Aetna, Cigna, UnitedHealthcare, CMS LCDs); charts CC0, policies quoted with their sources","held_out":true,"caveats":["Synthetic charts; real charts are longer and messier.","Each set was run once; fixes were made after each and measured on the next, fresh set. The latest set is small (16).","Requirement status right 91% to 94%, below the 95% target; 10% to 19% of requirements not found."],"date":"2026-09-28","doc_url":"https://decosa.ai/metrics/evals/prior-auth-check"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"decision and criteria accuracy","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"latest fresh held-out set (16 synthetic charts, 4 real policies, run once after the cold-user fixes): decision right","value":"14/16; unsupported requests called ready 1/11; supported called ready 4/5","source":"decosa-api docs/evals/prior-auth-check.md, measured on our server 2026-09-28, gateway route"},{"metric":"earlier fresh set (35 charts, 7 policies, after the guards) / first set (60 charts, 12 policies, before them)","value":"24/35 (unsupported called ready 4/21) / 45/60 (10/31)","source":"decosa-api docs/evals/prior-auth-check.md"},{"metric":"requirement status right (of requirements the model found)","value":"91.1%, 91.8%, 93.6%; 81% to 90% of gold requirements found","source":"decosa-api docs/evals/prior-auth-check.md"},{"metric":"undocumented read as not met (the appeal engine's main error, 6 of 80 cases)","value":"0 of 111 cases","source":"decosa-api docs/evals/prior-auth-check.md"},{"metric":"letter sentences rated unsupported by a blind reviewer","value":"6/375, 2/91, 0/46","source":"decosa-api docs/evals/prior-auth-check.md"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"decision and criteria accuracy","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"decision and criteria accuracy","value":"not measured yet","source":null}]}],"benchmark":{"title":"Does it say \"don't send\" when it should?","intro":"Synthetic charts written blind against 23 real published payer policies and sections (Aetna, Cigna, UnitedHealthcare, CMS LCDs), each labelled twice (author and blind reviewer agreed on all 744 requirements). Prompts were written on 4 other policies. Three sets were each run once: before the guards (60 cases), after them (35), and after fixes from two blind coordinator tests (16).","rows":[{"label":"Decision right, latest set","value":"14 of 16","detail":"earlier sets: 24 of 35, 45 of 60"},{"label":"Unsupported requests it called ready","value":"1 of 11","detail":"earlier sets: 4 of 21, 10 of 31"},{"label":"Supported requests it called ready","value":"4 of 5","detail":"earlier sets: 8 of 14, 26 of 29"},{"label":"Undocumented read as not met","value":"0 of 111 cases","detail":"the appeal engine's main error was 6 of 80"},{"label":"Cost per check","value":"about $0.014","detail":"median on the latest set, gateway list price; p95 $0.027"}],"points":[{"heading":"Where it fails","text":"It misses requirements hidden in footnotes, appendices and tables (13% to 19% aren't found), misreads a table lookup (a resection-weight scale by body surface area), and still calls a few unsupported requests ready. Treat ready as \"nothing obviously missing\", not a guarantee."},{"heading":"What it does not show","text":"Synthetic charts, one author per policy, no labels from working coordinators. Real charts are longer and messier. Check it on your own recent requests before relying on it."}],"source":"decosa-api docs/evals/prior-auth-check.md, 28 Sep 2026"},"verification":{"hosted":{"date":"2026-09-28","result":"partial","p50_ms":41300,"p95_ms":105500,"runs":16,"receipts_per_run":null,"cost_per_run_usd":0.0136},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, run with a named data volume, direct route to the local Qwen3.8-27B, local signing; torn down after","notes":"The rehearsal bundle passed 8/8 (CPAP ready with a letter citing the AHI 26.5, CPAP not supported with no letter, the not-met criterion quoting 3.6, the record verifies and fails once changed, every receipt attested) in 28.2 s; the lumbar MRI sample came back ready in 17.5 s and the InterQual sample can't-check in 2.9 s, all receipts attested. Model-server startup was not re-run."},"known_limits":["Time and cost from the latest held-out run (4 checks at a time on the shared gateway); replaced by production measurements after launch.","Accuracy is below the target set for this tool (95% of requirements right): a person checks every criterion before sending.","Synthetic charts written against real policy excerpts; not measured on real charts or with coordinators' labels.","Licensed criteria (InterQual, MCG) are out of scope; a policy that points to them gets can't check."],"receipt_coverage":"full"},"cost_per_run_usd":0.0136,"rehearsal_bundle":{"url":"/samples/prior-auth-check.zip","checks":8,"bytes":6348},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Picks the criteria section of a long policy, splits the policy into requirements, alternatives and exclusions, checks each against the chart, re-checks every not-met answer (and every exclusion answered met), answers the payer's form questions, drafts the letter of medical necessity when the chart supports every criterion, and judges every letter sentence (the grounding judge)","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/prior-auth-check","page":"/clinics/prior-auth-check","json":"/use-cases/prior-auth-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"jottings-note","num":"103","name":"Notes from your own jottings","status":"live","industries":["healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Risk and safety statements carried, fresh blind split #3 (verbatim rule)","value":"113 / 117","unit":null,"n":117,"split":"heldout","note":"40 sessions written blind, 110 planted statements, 15 decoys. 0 dropped, 4 changed (fixed after, not re-measured), 0 false alarms, 0 drafts blocked; 100 of 395 kept sentences replaced by the jotting word for word."},{"name":"Draft units not supported, fresh held-out split (29 Sep, blind review)","value":"7 / 805","unit":null,"n":805,"split":"heldout","note":"0.87%: 0 contradicted, 0 added clinical claims, 7 of 60 notes (3 medication changes the client reported written as fact). The 0.5% target was not met."},{"name":"Therapy practice drafts not supported (blind review, 70 notes)","value":"13 / 1,292","unit":null,"n":1292,"split":"synthetic","note":"29 / 1,283 before these fixes; 68 of 68 risk statements carried (before: one sentence dropped a written 'no SI')."},{"name":"Cost per note at list price (mean), held-out split","value":"0.0072","unit":"USD","n":60,"split":"heldout","note":"23 model calls: each kept sentence now gets a grounding check and a meaning check (was $0.0051, 15 calls)."},{"name":"Draft units not supported by the jottings (blind review)","value":"25 / 701","unit":null,"n":701,"split":"test","note":"16 unsupported, 9 contradicted; 5 flagged as an added clinical claim. 18 of 48 notes had at least one."},{"name":"Same, one plain prompt to the same model (baseline)","value":"751 / 1574","unit":null,"n":1574,"split":"test","note":"454 flagged as an added clinical claim; 48 of 48 notes had at least one."},{"name":"Draft units not supported, after the fixes (blind review)","value":"5 / 326","unit":null,"n":326,"split":"heldout","note":"24 new sessions written after the test run; 3 of 24 notes had at least one."},{"name":"Audited elements marked right","value":"379 / 384","unit":null,"n":384,"split":"test","note":"Eight elements per session: date, start, stop, modality, interventions, goals, response, plan."},{"name":"Missing elements caught","value":"56 / 57","unit":null,"n":57,"split":"test","note":null},{"name":"Handwritten lines read right","value":"99 / 99","unit":null,"n":99,"split":"test","note":"12 synthetic photos rendered with handwriting fonts; not real handwriting."}],"dataset":"Synthetic post-session jottings written by writer agents from a trap spec: 14 dev, 48 test (12 as rendered handwriting photos), 24 fresh sessions written after the test run (6 photos); gold element labels by the writers. On 29 Sep a second fresh split of 60 sessions (92 planted risk statements, 22 decoys) was written blind before the risk rule was tuned.","held_out":true,"caveats":["Synthetic sessions written by agents from a spec the builder wrote; rendered handwriting fonts, not real handwriting.","The unsupported-sentence numbers come from one blind reviewer model (Claude Opus); no licensed therapist has read the outputs.","The fresh split measured the fixes the test found; later fixes (from the fresh split and a cold-user test) are checked only on dev and unit tests.","Test risk-mention gold counted jokes as risk; the tool no longer does (the blind review called joke-based risk lines added claims).","Latency was measured on a shared gateway under load from other workloads.","The 29 Sep fixes that came from reading the second fresh split's errors are checked only on dev data and a redraft of those sessions.","The verbatim rule's last fix (quote the whole risk jotting) came from reading split #3's errors; a fourth blind split is needed to measure it."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/jottings-note"},"quality_evidence":[{"tier":"standard","label":"Standard · one GPU for the model (hosted demo)","evidence":[{"metric":"Draft sentences and header lines a blind reviewer found not supported by the jottings (48 held-out synthetic sessions)","value":"25 / 701","source":"docs/evals/jottings-note.md, test split, 2026-09-28 (blind Claude Code Opus review)"},{"metric":"The same blind check on 24 new sessions after the fixes the test found","value":"5 / 326","source":"docs/evals/jottings-note.md, fresh split, 2026-09-28"},{"metric":"Same reviewer, the same sessions drafted by one plain prompt to the same model (no checks)","value":"751 / 1574","source":"docs/evals/jottings-note.md, test split"},{"metric":"Audited elements marked present or missing correctly","value":"379 / 384","source":"docs/evals/jottings-note.md, test split"},{"metric":"Elements missing from the jottings that it marked missing","value":"56 / 57","source":"docs/evals/jottings-note.md, test split"}]},{"tier":"best","label":"Best · adds the second reader and dictation","evidence":[{"metric":"Handwritten lines read right (12 rendered photos, test split)","value":"99 / 99","source":"docs/evals/jottings-note.md, test split"},{"metric":"Lines flagged 'check the reading' that were in fact read right","value":"16 of 16 flags","source":"docs/evals/jottings-note.md, test split (the second reader's disagreements are mostly punctuation or letter case)"},{"metric":"Dictations of up to 60 s transcribed and drafted; a 190 s recording refused","value":"4 / 4; refused","source":"docs/evals/jottings-note.md, self-host check (synthetic voice)"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":25462,"p95_ms":30257,"runs":5,"receipts_per_run":24,"cost_per_run_usd":0.008581},"selfhost":{"date":"2026-09-28","result":"pass","method":"fresh clone of decosa-api on our server, api image built from docker/api/Dockerfile, compose api with a named data volume on the host network, direct route to the running Qwen3.8-27B, page parser and ASR; local signing; torn down after","notes":"Rehearsal bundle 12/12 four times in a row (8.5-9.6 s); smoke ok in 6.2 s with 17 signed receipts; the photo sample read 9 lines (8 agreed by both readers); 4 synthetic dictations (computer voice, 17-32 s) transcribed and drafted, a 190 s recording refused; /jottings/app drafted a note at 1280 and 390 px with no horizontal scroll; a cold-user test ran a 20-session catch-up against it. Model-server startup was not re-run."},"known_limits":["Hosted verification on production (decosa.ai, gateway route), 28 Sep 2026: the smoke (cbt-panic sample) run 5 times one at a time.","Measured on synthetic sessions written by agents and on rendered handwriting; no real jottings, real handwriting or licensed therapist's review yet.","Not zero: on 24 new sessions 5 of 326 draft units were still not supported by the jottings (mostly shorthand read the wrong way, such as 'sat' for Saturday).","The same jottings can give different drafts on two runs.","Dictation was tested with a computer voice only; spoken dates and times are converted to digits, other numbers stay as words.","Needs a 32-96 GB GPU in the practice for real notes until a confidential hosted tier with a BAA exists."],"receipt_coverage":"full"},"cost_per_run_usd":0.008581,"rehearsal_bundle":{"url":"/samples/jottings-note.zip","checks":12,"bytes":3093},"models":[{"name":"decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)","role":"Reading typed jottings, the completeness check (dates, times and risk in code), the number, clinical-claim, qualifier and attribution guards, layouts, catch-up and the signed record (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Drafts the sentences, tags the audited elements, checks every sentence (the grounding judge), rewrites a failed sentence once, and reads a photo's page","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/jottings-note","page":"/clinics/jottings-note","json":"/use-cases/jottings-note.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"evidence-runner","num":"140","name":"Capture audit evidence from your admin screens","status":"preview","industries":["compliance-trust","software"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Quarterly verdicts right on held-out consoles (no model)","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Four held-out made-up consoles; 9 / 9 changed screens caught, 0 false alarms."},{"name":"Quarterly verdicts right on development consoles","value":"19 / 19","unit":null,"n":19,"split":"dev","note":"7 / 7 changed screens caught."},{"name":"Setup: screen found and every named setting read right, held-out after fixes","value":"25 / 26","unit":null,"n":26,"split":"test","note":"Three held-out consoles on the hosted gateway, after four bugs found on the first look were fixed (a second look)."},{"name":"Setup, first look at two held-out consoles","value":"10 / 17","unit":null,"n":17,"split":"test","note":"Before the fixes; the failures exposed four bugs."},{"name":"Setup, first look at a console added after every fix","value":"8 / 10","unit":null,"n":10,"split":"test","note":"Both misses came from one engine bug (a menu link read as an action), fixed after this run."},{"name":"Setup, development consoles","value":"19 / 20","unit":null,"n":20,"split":"dev","note":null},{"name":"Wrong values read at setup","value":"0","unit":null,"n":null,"split":"test","note":"In every run; values code could not place were left for the person."},{"name":"Redesigned console handled (routes repaired)","value":"19 / 19","unit":null,"n":19,"split":"dev","note":null},{"name":"Writes that reached a console","value":"0","unit":null,"n":null,"split":"test","note":"Every run, including pages that asked agents to reset and save."}],"dataset":"Six made-up admin consoles written for this eval (identity admin, cloud console, an own-product admin with iframe pages, an endpoint manager with dialogs, code-hosting organisation settings, an HR and payroll system), two quarters of settings each, a redesigned variant of the two development consoles, planted text aimed at AI agents and session expiries.","held_out":false,"caveats":["The same author wrote the consoles and the runner; the consoles copy real shapes but are not real products.","Held-out consoles were looked at more than once: the first look exposed bugs that were fixed, so later numbers on them are second looks. The first-look numbers are listed separately.","No real tenant, no human reviewer timing; review time is not measured.","Small n per console (7-10 screens)."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/evidence-runner"},"quality_evidence":[{"tier":"lite","label":"Lite · quarterly runs only, no GPU","evidence":[{"metric":"Quarterly verdicts right, six made-up consoles","value":"52 / 52 screens (33 / 33 on held-out consoles); 16 / 16 changed screens caught; 0 false alarms","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"}]},{"tier":"standard","label":"Standard · setup and repairs with Qwen3.8-27B","evidence":[{"metric":"Setup, held-out consoles after fixes","value":"25 / 26 screens found with every named setting read right; 0 wrong values","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026 (hosted gateway run)"},{"metric":"Setup, first look at a console added after every fix","value":"8 / 10 (both misses from one bug, fixed after)","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"},{"metric":"Redesigned console, routes repaired","value":"19 / 19 screens","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":7500,"p95_ms":10200,"runs":3,"receipts_per_run":0,"cost_per_run_usd":0},"selfhost":{"date":"2026-09-29","result":"pass","method":"fresh clone, venv, tests, then the CLI attached over CDP to a separate Chromium profile standing in for the admin's own browser, against a made-up console served over plain HTTP","notes":"Setup 3 of 3 screens right in 46 s on the local model server; next quarter 2 changes found in 2 s (exit 1), 0 writes reached the console, the stand-in browser's own tab left untouched."},"known_limits":["Made-up consoles only so far; no real tenant has been run.","Only the settings you name are asserted in the certificate; other settings on the same screen are compared and reported, not asserted.","The URL and time band is drawn onto each PNG by the runner (the certificate binds the file to its capture); it is not the browser's own address bar or the system clock.","Microsoft 365 and Entra admin centers: witnessed capture only (no agent navigation); use Microsoft Graph for settings.","Setup needs a person to approve the baseline values and to handle screens left for a person (pages with text aimed at AI agents, screens the agent could not find).","Hosted runs are demos on made-up consoles; real consoles run on your side."],"receipt_coverage":"partial"},"cost_per_run_usd":0,"rehearsal_bundle":{"url":"/samples/evidence-runner.zip","checks":10,"bytes":1245},"models":[{"name":"decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)","role":"Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B","role":"Setup and repairs only: finds each named screen (read-only) and points at rows when code cannot place a setting","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/evidence-runner","page":"/tools/finance/evidence-runner","json":"/use-cases/evidence-runner.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"ehr-drafts","num":"145","name":"Put the visit into your EHR as drafts","status":"preview","industries":["healthcare"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Typed or picked values wrong (read back from the EHR's database)","value":"0 / 419","unit":null,"n":419,"split":"synthetic","note":"95% CI 0-0.91%. 32 synthetic visits, clean run 2 (quantity and refills only when said). Run 1: 0 / 534."},{"name":"Items entered, of those the rules allow the agent to enter","value":"101 / 104","unit":null,"n":104,"split":"synthetic","note":"Misses: 2 inhalers wrongly held by the medicine check, 1 lab order the agent could not finish. Left for you by rule: 14 prescriptions with no quantity said, 2 controlled substances, stops and follow-ups."},{"name":"Wrong-patient attempts that wrote to another chart","value":"0 / 20","unit":null,"n":20,"split":"synthetic","note":"16 wrong charts open at the start (near-duplicate name, date of birth or MRN; another patient): stopped with 0 entries. 4 chart switches mid-run: caught before the next save."},{"name":"Forced sign, transmit, send, fax, e-mail, bill and delete requests held","value":"80 / 80","unit":null,"n":80,"split":"synthetic","note":"8 kinds x 10 visits, sent from inside the page after the drafts; plus 16 of 16 clicks on the app's own eSign, Transmit Order and Fax buttons held."},{"name":"Planted medicine errors flagged by the medicine check","value":"90 / 90","unit":null,"n":90,"split":"synthetic","note":"Dose, frequency, drug (sound-alikes) and duration; sound-alike pairs written after the fix: 28 / 28. End to end: 8 of 8 never entered."},{"name":"Planted instructions in the chart that caused harm","value":"0 / 10","unit":null,"n":10,"split":"synthetic","note":"Encounter reason or an intake note, English and Spanish; 9 of 10 runs paused on the text."},{"name":"Correct medicines wrongly held by the medicine check","value":"2 / 28","unit":null,"n":28,"split":"synthetic","note":"Both an inhaler with a spoken route."}],"dataset":"32 synthetic primary-care visits (18 templates, made-up patients): items fixed in code as the ground truth, transcripts written by Qwen3.8-27B around them, on a self-hosted OpenEMR 7.0.3 test instance.","held_out":false,"caveats":["Synthetic visits and transcripts; the model that wrote the transcripts is the model that drives the agent.","One EHR (OpenEMR). The agent's instructions per form (the EHR profile) were written while testing on the same forms.","The medicine checker was trained on Qwen-written visits; its numbers on these transcripts are likely optimistic.","Values were checked against the visit output, not against what a clinician would have wanted: a wrong visit output would be copied faithfully.","Transcripts were patched twice before the measured runs, both recorded: amounts per dose, and quantities and refills said in half the visits.","Two blind cold users (a solo physician: maybe / maybe; a practice manager: no / no) read a one-page description; neither used it on real work."],"date":"2026-09-29","doc_url":null},"quality_evidence":[{"tier":"standard","label":"Standard · one GPU for the model","evidence":[{"metric":"Typed and picked values wrong, read back from the EHR's database (32 synthetic visits)","value":"0 / 419","source":"decosa-api docs/evals/ehr-drafts.md, clean run 2, 2026-09-29 (run 1: 0 / 534)"},{"metric":"Wrong-patient attempts that wrote to another chart","value":"0 / 20","source":"decosa-api docs/evals/ehr-drafts.md, 2026-09-29"},{"metric":"Planted wrong doses, drugs, frequencies and durations flagged by the medicine check","value":"90 / 90","source":"decosa-api docs/evals/ehr-drafts.md, 2026-09-29"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":90300,"p95_ms":162900,"runs":32,"receipts_per_run":null,"cost_per_run_usd":0.0341},"selfhost":null,"known_limits":["Measured on one EHR only: self-hosted OpenEMR 7.0.3 with made-up patients. A second open-source EHR (OpenMRS O3) was tried and did not start cleanly; not measured.","Not yet packaged as the browser extension; the engine ran in a test browser against the test EHR.","A prescription whose quantity was not said goes to your list (OpenEMR requires a quantity, and the agent never guesses one).","The medicine check wrongly held 2 of 28 correct medicines (an inhaler whose route was said).","The entries are saved under your login, so the EHR's audit log shows you; each draft's comment or note says it came from ehr-drafts, and the signed record shows every step.","About 1.5 to 3 minutes per visit on a shared model; no batch mode yet."],"receipt_coverage":"full"},"cost_per_run_usd":0.0341,"rehearsal_bundle":null,"models":[{"name":"Qwen3.8-27B on the Decosa computer-use engine","role":"The agent that fills the EHR's forms: one receipted decision per step, values only from the visit, the one draft save released after the chart and form are checked in code","license":"Apache-2.0","hf_repo":"Qwen/Qwen3.8-27B-FP8"},{"name":"decosa-note-detail-checker (M17)","role":"Reads each medicine (drug, dose, frequency, route when said, duration) against the transcript lines it came from","license":"Apache-2.0","hf_repo":"decosaai/decosa-note-detail-checker-modernbert-large"},{"name":"decosa-api ehr-drafts and the computer-use network gate","role":"Holds sign, finalise, transmit, send, fax, e-prescribe, bill and delete requests in the browser; releases each draft save once","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/ehr-drafts","page":"/clinics/ehr-drafts","json":"/use-cases/ehr-drafts.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"demand-reader","num":"155","name":"Injury demand reader","status":"preview","industries":["finance","legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Respond-by date exact","value":"7 of 7","unit":null,"n":7,"split":"test","note":"every time-limited letter in the test set; the one ordinary demand was read as not time-limited (8 of 8)"},{"name":"Statute's minimum exact","value":"7 of 7","unit":null,"n":7,"split":"test","note":null},{"name":"Conditions of acceptance found","value":"31 of 31","unit":null,"n":31,"split":"test","note":"precision 33 of 41"},{"name":"Billed totals within $1","value":"8 of 8","unit":null,"n":8,"split":"test","note":"0 invented bill lines"},{"name":"Value or payment language in the tool's own words","value":"0","unit":null,"n":8,"split":"test","note":null}],"dataset":"24 synthetic injury demand packages (letters, bills and records; 406 pages; CA, GA, MO, MT, UT, FL, TX, OH) written blind by a separate author with planted deadlines, conditions, missing elements, specials gaps and citation problems; dev 16, test 8.","held_out":true,"caveats":["Synthetic packages from one author; real ones are longer, scanned and messier.","Small test set (8 packages): wide confidence intervals.","Prompts and rules were tuned on the 16 dev packages.","Six states encoded; the rest are not."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/demand-reader"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"respond-by date exact, held-out test (8 packages, gateway, run once)","value":"7 of 7","source":"decosa-api docs/evals/demand-reader.md, measured on our server 2026-09-29, gateway route"},{"metric":"conditions of acceptance found / billed totals within $1 (test)","value":"31 of 31 / 8 of 8","source":"decosa-api docs/evals/demand-reader.md, measured on our server 2026-09-29, gateway route"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":161000,"p95_ms":208000,"runs":8,"receipts_per_run":23,"cost_per_run_usd":0.0197},"selfhost":null,"known_limits":["Hosted verification ran on our pre-release server through the production gateway, before these routes reached the production API.","Measured on 24 synthetic packages written by one author; real packages are longer, scanned and messier.","Typed or pasted page text only in this version; scans go through the document reader first.","Six states encoded; the rest return 'not encoded: check with counsel'."],"receipt_coverage":"full"},"cost_per_run_usd":0.0197,"rehearsal_bundle":{"url":"/samples/demand-reader.zip","checks":8,"bytes":3612},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the letter's terms and every condition with quotes, each bill page's charge lines, and each record page's visits (the medical chronology's page extractor); the grounding judge checks every sentence it writes","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/demand-reader","page":"/tools/insurance/demand-reader","json":"/use-cases/demand-reader.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"certificate-check","num":"156","name":"Certificate request check","status":"preview","industries":["finance"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Status right on found requirements","value":"109 of 120 (0.91)","unit":null,"n":120,"split":"test","note":"third test run after cold-user fixes (runs: 0.91, 0.84, 0.91); CI 0.88-0.95"},{"name":"Requirements found","value":"120 of 124 (0.97)","unit":null,"n":124,"split":"test","note":"precision 120 of 144"},{"name":"'Met' calls that were right","value":"102 of 103","unit":null,"n":103,"split":"test","note":"a wrong 'met' is the E&O risk"},{"name":"'Met' calls without a policy quote","value":"0","unit":null,"n":103,"split":"test","note":null}],"dataset":"27 synthetic contract, request and policy triples (354 labelled requirements, 17 states), written blind by a separate author; stratified split, dev 18, test 9.","held_out":false,"caveats":["Synthetic papers from one author; 78% of labels are 'met', so read the per-status numbers.","The test set was run three times; the third run followed a fix prompted by the second, so it is not fully held out.","Not a coverage opinion; it compares wording.","Slow under load: about two minutes a request on the shared gateway."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/certificate-check"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"status right on found requirements, test set (9 triples, 120 requirements, gateway; third run, not fully held out)","value":"109 of 120 (0.91)","source":"decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route"},{"metric":"requirements found / 'met' calls that were right (test, third run)","value":"120 of 124 / 102 of 103","source":"decosa-api docs/evals/certificate-check.md, measured on our server 2026-09-29, gateway route"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":33257,"p95_ms":33845,"runs":5,"receipts_per_run":31,"cost_per_run_usd":0.014522},"selfhost":null,"known_limits":["Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","Measured on 27 synthetic contract, request and policy triples written by one author; real policies run to 100+ pages of forms.","One wrong 'met' on the test set (an umbrella that excludes pollution); blanket waivers of subrogation often come back as needing a person.","Typed or pasted text only; scanned policies need the document reader first."],"receipt_coverage":"full"},"cost_per_run_usd":0.014522,"rehearsal_bundle":{"url":"/samples/certificate-check.zip","checks":5,"bytes":3964},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Lists each requirement from the contract and the email with a verbatim quote, then per requirement says whether the policy papers meet it, quoting the policy; the grounding judge checks each reason","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/certificate-check","page":"/tools/insurance/certificate-check","json":"/use-cases/certificate-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"bank-change-check","num":"157","name":"Vendor bank-change check","status":"preview","industries":["finance","compliance-trust"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Fraud flagged (score 3 or more)","value":"9 of 9","unit":null,"n":9,"split":"test","note":"Wilson 95% CI 0.70-1.00; dev 21 of 21"},{"name":"Genuine changes flagged","value":"0 of 8","unit":null,"n":8,"split":"test","note":"dev 0 of 18"},{"name":"Bank-change requests detected","value":"19 of 19","unit":null,"n":19,"split":"test","note":"dev 45 of 45"},{"name":"Content signs found: open model vs keyword rules","value":"43 of 47 vs 23 of 47","unit":null,"n":47,"split":"synthetic","note":"dev and test together; precision 0.83 vs 0.88"},{"name":"Verdicts calling an email safe","value":"0","unit":null,"n":64,"split":"synthetic","note":null}],"dataset":"64 synthetic emails (30 fraud, 26 genuine, 8 with no bank change; 6 not in English) and a 26-vendor file, written blind by a separate author; stratified split, dev 45, test 19.","held_out":true,"caveats":["Synthetic emails from one author; real BEC mail is messier.","Small test set (9 frauds, 8 genuine changes): wide confidence intervals.","Code signs were tuned on the dev set.","A hacked real mailbox writing calmly passes every check in the email; the call-back is the control."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/bank-change-check"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"fraud flagged / genuine changes flagged, held-out test (19 emails, gateway, run once)","value":"9 of 9 / 0 of 8","source":"decosa-api docs/evals/bank-change-check.md, measured on our server 2026-09-29, gateway route"},{"metric":"content signs found, open model vs keyword rules (all 64 emails)","value":"43 of 47 vs 23 of 47","source":"decosa-api docs/evals/bank-change-check.md, dev and test together"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":8100,"p95_ms":18900,"runs":19,"receipts_per_run":1,"cost_per_run_usd":0.00053},"selfhost":null,"known_limits":["Hosted verification ran on our pre-release server through the production gateway, before these routes reached the production API.","Measured on 64 synthetic emails written by one author; real business email compromise is messier.","A calm email from a real vendor's hacked mailbox that keeps the same bank country passes every check in the email; only the call-back catches it.","No domain age, ownership or account-validation look-ups (nothing leaves the box by design)."],"receipt_coverage":"full"},"cost_per_run_usd":0.00053,"rehearsal_bundle":{"url":"/samples/bank-change-check.zip","checks":8,"bytes":4471},"models":[{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Reads the email's text and quotes the pressure, secrecy, 'don't call' and redirected-payment signs, and reads the new account's bank and holder; the header, domain and vendor-file checks are code","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/bank-change-check","page":"/tools/finance/bank-change-check","json":"/use-cases/bank-change-check.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"family-film","num":"170","name":"Family interview film","status":"live","industries":["creative-media","entertainment"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Shown quotes inside her own turn, at the right time","value":"46 / 46","unit":null,"n":46,"split":"dev","note":"5 languages; a quote counts when it overlaps her true turn by at least 80% of its span."},{"name":"Shown quotes that match her words (fuzzy 0.9)","value":"44 / 46","unit":null,"n":46,"split":"dev","note":"Both misses write the year as digits where the script spells it out; one also has \"né\" for \"née\" (a real slip)."},{"name":"Quotes dropped by the re-hearing check","value":"2 / 48","unit":null,"n":48,"split":"dev","note":"Both had speech-recognition slips; dropped quotes are never shown."},{"name":"Speaker labels right","value":"79 / 80","unit":null,"n":80,"split":"dev","note":"The miss: a 0.66 s interviewer segment labelled hers; no quote came from it."},{"name":"Meaning check: lines flagged / clear mistranslations","value":"27 / 46 flagged; 2 clear","unit":null,"n":46,"split":"dev","note":"The builder read every flag: 2 clear mistranslations, 25 nitpicks or misreadings. Misses not measured."},{"name":"Consent refusals and hesitations refused","value":"2 / 2","unit":null,"n":2,"split":"synthetic","note":null},{"name":"Photo backs read","value":"3 / 3","unit":null,"n":3,"split":"synthetic","note":"Synthetic handwriting on the sample's photo backs."},{"name":"First trailer after upload (p50)","value":"40.8 s","unit":null,"n":5,"split":"dev","note":"37-71 s; interviews of 56-207 s, pre-release server."},{"name":"Cost per film","value":"$0.0085-0.02","unit":null,"n":5,"split":"dev","note":"List prices for the text model and GPU time."},{"name":"Blind granddaughter: keeps it / shares it as is / would pay","value":"yes / no / $39","unit":null,"n":1,"split":"synthetic","note":"An Opus sub-agent on the Italian sample: about 40 min to fix in the tool against 6+ hours by hand (its estimate). Two of its problems fixed since: held error lines, no borrowed photos."}],"dataset":"5 synthetic interviews written by the builder (Italian 207 s, Spanish 70 s, Portuguese 68 s, French 56 s, German 81 s), voiced by VoxCPM2 voice design (Apache-2.0; no real person recorded or cloned), with public-domain Library of Congress photos and synthetic photo backs.","held_out":false,"caveats":["Synthetic interviews with two clean voices; real recordings are not measured.","The same builder wrote the scripts, the pipeline and the eval, and fixed bugs found in an earlier run on the same set.","Quote correctness is a fuzzy text match against the script, so years written as digits count as misses.","The meaning-check judgement is the builder's own reading, not a blind or native-speaker review."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/family-film"},"quality_evidence":[{"tier":"lite","label":"Lite · chapters, quotes, subtitles and the film, no photo reading","evidence":[{"metric":"Shown quotes inside her own turn, at the right time (5 synthetic interviews)","value":"46 / 46","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Speaker labels right","value":"79 / 80 segments","source":"decosa-api docs/evals/family-film.md, 2026-09-29"}]},{"tier":"standard","label":"Standard · the hosted demo, with photo backs read","evidence":[{"metric":"Shown quotes inside her own turn, at the right time (5 synthetic interviews)","value":"46 / 46","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Photo backs read (synthetic handwriting)","value":"3 / 3","source":"decosa-api docs/evals/family-film.md, 2026-09-29"},{"metric":"Meaning check: lines flagged / clear mistranslations among them","value":"27 of 46 flagged; 2 clear mistranslations","source":"decosa-api docs/evals/family-film.md, 2026-09-29; the builder read every flag"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":40840,"p95_ms":70540,"runs":5,"receipts_per_run":11,"cost_per_run_usd":0.0092},"selfhost":null,"known_limits":["Measured on 5 synthetic interviews with two voices each, voiced by an open voice-design model; real family recordings (noise, overlapping talk, more relatives) are not measured.","The meaning check flags about half the lines, and most flags are nitpicks: read the English yourself for a language you know.","Quotes copy the speech recogniser's words: a one-letter slip (\"né\" for \"née\") reached a shown quote.","Chapters follow the order she told them, not the years, and can't be reordered yet.","One film style."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0092,"rehearsal_bundle":{"url":"/samples/family-film.zip","checks":8,"bytes":104263},"models":[{"name":"Qwen3-ASR-1.7B (language pack speech service)","role":"Speech recognition in her language, one call per speech segment (so every word keeps its time)","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-ASR-1.7B"},{"name":"ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)","role":"Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hers","license":"Apache-2.0","hf_repo":"speechbrain/spkrec-ecapa-voxceleb"},{"name":"Qwen3.8-27B (NVFP4)","role":"Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallback","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Hy-MT2-7B (the language-pack block)","role":"English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every line","license":"Apache-2.0","hf_repo":"tencent/Hy-MT2-7B"},{"name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","role":"Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward faces","license":"MIT","hf_repo":null},{"name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","role":"Reads the handwriting on the backs of photos (place, year, names) to date and place each picture","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"},{"name":"ACE-Step 1.5 cue (music-gen-cleared library)","role":"A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)","license":"MIT","hf_repo":"ACE-Step/Ace-Step1.5"},{"name":"decosa-api family_film module + FFmpeg + c2pa-python","role":"The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per file","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/family-film","page":"/apps/family-film","json":"/use-cases/family-film.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"music-video-starring-you","num":"171","name":"Music video starring you","status":"preview","industries":["music","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planned cuts found on the beat in the exports","value":"27 / 27","unit":null,"n":27,"split":"synthetic","note":"3 storyboard renders of the sample song; median 5.7 ms, max 16.3 ms from the beat grid (half a frame at 30 fps is 16.7 ms)."},{"name":"Storyboard frames flagged by the frame safety check","value":"0 / 30","unit":null,"n":30,"split":"synthetic","note":"Sample performers are synthetic adults; the check fails closed."},{"name":"Storyboard video done (p50)","value":"88.2 s","unit":null,"n":3,"split":"synthetic","note":"86-105 s (p50 88.2 s); first picture drawn after 12-20 s; shared GPU."},{"name":"Cost per storyboard video","value":"$0.029-0.034","unit":null,"n":3,"split":"synthetic","note":"Text model at list price plus GPU time at $1.32/h."},{"name":"MiniMax H3 time per moving shot","value":"57-65 s (one reference)","unit":null,"n":10,"split":"synthetic","note":"Sketch tier 864x480, fp8, one 96 GB card, peak 49.6-52 GiB; recorded in a GPU window on 29 Sep."},{"name":"MiniMax H3 cost per shot","value":"$0.022-0.029","unit":null,"n":27,"split":"synthetic","note":"GPU time at $1.32/h; 27 shots including 6 re-rolls."},{"name":"Face likeness to the consent clip (ArcFace cosine, storyboard / H3)","value":"0.52 / 0.36","unit":null,"n":null,"split":"synthetic","note":"Internal QC with a non-commercial model, not shipped. H3 shots were often wide or turned away (faces found in 15 of 80 sampled frames); frontal re-rolls reached 0.51."},{"name":"Moving shots that kept the face (likeness gate)","value":"5 / 10","unit":null,"n":10,"split":"synthetic","note":"One middle frame per moving shot checked against each person's reference; the rest fell back to the still. Some wrong faces still pass."},{"name":"Blind artist: posts it / would pay (storyboard; moving)","value":"yes, $20; no, $0","unit":null,"n":1,"split":"synthetic","note":"An Opus sub-agent as an independent artist: the storyboard as a teaser and Canvas, not as the music video; the moving shots lost her face too often. Before the likeness gate on moving shots."}],"dataset":"Synthetic sample performers (consent clips made from designed faces and voices; no real person), one sample song made with MiniMax-Music3, 3 storyboard renders per mode, and 27 MiniMax H3 shots (10 artist, 11 couple, 6 re-rolls) rendered in a GPU window on 29 Sep 2026.","held_out":false,"caveats":["Synthetic performers and one song; real selfies and real tracks are not measured.","Beat offsets are against the detected beat grid, not a human one.","Visual quality is not scored; some H3 sketch shots show artefacts (gold squiggles in one look, a banding glitch).","ArcFace (buffalo_l) is non-commercial and used only as internal QC.","The same builder wrote the pipeline and the eval."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/music-video-starring-you"},"quality_evidence":[{"tier":"standard","label":"Standard · the hosted demo, a storyboard cut on the beat","evidence":[{"metric":"Planned cuts found on the beat (3 renders)","value":"27 / 27; median 5.7 ms, max 16.3 ms","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Frames flagged by the frame safety check","value":"0 / 30","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Visual quality","value":"not scored; drawn stills with camera moves","source":null}]},{"tier":"best","label":"Best · moving shots on MiniMax H3, one 96 GB card","evidence":[{"metric":"Time per moving shot (sketch 864x480)","value":"57-65 s with one reference","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Face likeness to the consent clip (ArcFace, internal QC)","value":"0.36 (frontal re-rolls 0.51)","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29; faces often small or turned"},{"metric":"Blind testers who would pay for moving shots","value":"0 of 2 (both would for the storyboard)","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":88220,"p95_ms":104700,"runs":3,"receipts_per_run":5,"cost_per_run_usd":0.0291},"selfhost":null,"known_limits":["Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).","No lip-sync to the vocal.","Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.","The adult check is a vision estimate from the consent frames, not an ID check.","Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.","The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).","Measured with synthetic sample performers and one sample song."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0291,"rehearsal_bundle":{"url":"/samples/music-video-starring-you.zip","checks":10,"bytes":1342},"models":[{"name":"Qwen3-ASR-1.7B (language pack speech service)","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-ASR-1.7B"},{"name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","license":"MIT","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checks","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa-mvideo-analyze (services/mvideo)","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","license":"AGPL-3.0-or-later (decosa-api)","hf_repo":null},{"name":"FLUX.2 klein 4B","role":"Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)","license":"Apache-2.0","hf_repo":"black-forest-labs/FLUX.2-klein-4B"},{"name":"decosa-api starring module + FFmpeg + c2pa-python","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/music-video-starring-you","page":"/apps/music-video-starring-you","json":"/use-cases/music-video-starring-you.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"our-story-film","num":"172","name":"Our story film","status":"preview","industries":["creative-media","entertainment"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Planned cuts found on the beat in the exports","value":"30 / 30","unit":null,"n":30,"split":"synthetic","note":"3 storyboard renders of the sample song; median 5.2 ms, max 16.0 ms from the beat grid (half a frame at 30 fps is 16.7 ms)."},{"name":"Storyboard frames flagged by the frame safety check","value":"0 / 33","unit":null,"n":33,"split":"synthetic","note":"Sample performers are synthetic adults; the check fails closed."},{"name":"Storyboard video done (p50)","value":"142.5 s","unit":null,"n":3,"split":"synthetic","note":"122-275 s (p50 142.5 s); first picture drawn after 15-17 s; shared GPU."},{"name":"Cost per storyboard video","value":"$0.040-0.074","unit":null,"n":3,"split":"synthetic","note":"Text model at list price plus GPU time at $1.32/h."},{"name":"MiniMax H3 time per moving shot","value":"about 80 s (two references)","unit":null,"n":11,"split":"synthetic","note":"Sketch tier 864x480, fp8, one 96 GB card, peak 49.6-52 GiB; recorded in a GPU window on 29 Sep."},{"name":"MiniMax H3 cost per shot","value":"$0.022-0.029","unit":null,"n":27,"split":"synthetic","note":"GPU time at $1.32/h; 27 shots including 6 re-rolls."},{"name":"Face likeness to the consent clip (ArcFace cosine, storyboard / H3)","value":"0.43 / 0.40","unit":null,"n":null,"split":"synthetic","note":"Internal QC with a non-commercial model, not shipped. H3 shots were often wide or turned away (faces found in 15 of 80 sampled frames); frontal re-rolls reached 0.51."},{"name":"Moving shots that kept the face (likeness gate)","value":"5 / 11","unit":null,"n":11,"split":"synthetic","note":"One middle frame per moving shot checked against each person's reference; the rest fell back to the still. Some wrong faces still pass."},{"name":"Blind partner: gives it / would pay (storyboard; moving)","value":"yes after one fix, $29; no, $0","unit":null,"n":1,"split":"synthetic","note":"An Opus sub-agent buying an anniversary gift. The fix it asked for (his partner's face changed in one shot) is why both partners are now checked."}],"dataset":"Synthetic sample performers (consent clips made from designed faces and voices; no real person), one sample song made with MiniMax-Music3, 3 storyboard renders per mode, and 27 MiniMax H3 shots (10 artist, 11 couple, 6 re-rolls) rendered in a GPU window on 29 Sep 2026.","held_out":false,"caveats":["Synthetic performers and one song; real selfies and real tracks are not measured.","Beat offsets are against the detected beat grid, not a human one.","Visual quality is not scored; some H3 sketch shots show artefacts (gold squiggles in one look, a banding glitch).","ArcFace (buffalo_l) is non-commercial and used only as internal QC.","The same builder wrote the pipeline and the eval."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/music-video-starring-you"},"quality_evidence":[{"tier":"standard","label":"Standard · the hosted demo, a storyboard cut on the beat","evidence":[{"metric":"Planned cuts found on the beat (3 renders)","value":"30 / 30; median 5.2 ms, max 16.0 ms","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Frames flagged by the frame safety check","value":"0 / 33","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Visual quality","value":"not scored; drawn stills with camera moves","source":null}]},{"tier":"best","label":"Best · moving shots on MiniMax H3, one 96 GB card","evidence":[{"metric":"Time per moving shot (sketch 864x480)","value":"about 80 s with two references","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"},{"metric":"Face likeness to the consent clip (ArcFace, internal QC)","value":"0.40","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29; faces often small or turned"},{"metric":"Blind testers who would pay for moving shots","value":"0 of 2 (both would for the storyboard)","source":"decosa-api docs/evals/music-video-starring-you.md, 2026-09-29"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":142540,"p95_ms":274860,"runs":3,"receipts_per_run":5,"cost_per_run_usd":0.0405},"selfhost":null,"known_limits":["Moving shots need a render GPU, which is off until launch: the hosted demo makes a storyboard (drawn stills with camera moves).","No lip-sync to the vocal.","Likeness is checked by the text model on each frame; ArcFace numbers are internal QC only.","The adult check is a vision estimate from the consent frames, not an ID check.","Moving H3 shots aren't ready: blind testers rejected them (faces lost or someone else's); the likeness gate keeps about half and still lets some wrong faces through.","The likeness check is a yes/no from the text model on each picture; it caught 1 unlike picture in the couple test after the fix, and missed one before it (only the first partner was checked).","Measured with synthetic sample performers and one sample song."],"receipt_coverage":"partial"},"cost_per_run_usd":0.0405,"rehearsal_bundle":{"url":"/samples/our-story-film.zip","checks":10,"bytes":1412},"models":[{"name":"Qwen3-ASR-1.7B (language pack speech service)","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-ASR-1.7B"},{"name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","license":"MIT","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Adult check on the consent frames (vision), the shot list per song section from your three memories, and the frame safety and likeness checks","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"decosa-mvideo-analyze (services/mvideo)","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","license":"AGPL-3.0-or-later (decosa-api)","hf_repo":null},{"name":"FLUX.2 klein 4B","role":"Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)","license":"Apache-2.0","hf_repo":"black-forest-labs/FLUX.2-klein-4B"},{"name":"decosa-api starring module + FFmpeg + c2pa-python","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/our-story-film","page":"/apps/our-story-film","json":"/use-cases/our-story-film.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"settlement-video","num":"173","name":"Settlement video from the case file","status":"live","industries":["legal"],"deploy":["selfhost"],"eval_summary":{"metrics":[{"name":"Cites on the right page, rendered lines","value":"376 / 376","unit":null,"n":376,"split":"test","note":"6 held-out synthetic matters; prompts and checks frozen before"},{"name":"Dates in rendered lines right","value":"53 / 53","unit":null,"n":53,"split":"test","note":null},{"name":"Cites on the right page, matters written blind by another agent","value":"143 / 144","unit":null,"n":144,"split":"heldout","note":"139 / 144 by the writer's answer key; on review, four of the five unmatched cites were on the right page (findings the key did not list)"},{"name":"Cites on the right page, post-fix held-out matters","value":"201 / 201","unit":null,"n":201,"split":"test","note":"3 new seeds run once after two fixes made on the first held-out run"},{"name":"Every itemized charge read and summed to the cent","value":"11 / 11","unit":"matters","n":11,"split":"test","note":"the default total also leaves out charges with no matching visit; it equals the key in 7 / 11 because the chronology missed 1-3 visits in four matters"},{"name":"Planted bill problems flagged","value":"43 / 44","unit":null,"n":44,"split":"test","note":"before the injury 11/11, duplicate 11/11, no matching visit 11/11 (7 false flags), statement total off 10/11 (one total on a scan was unreadable and said so)"},{"name":"Rendered lines a blind judge found supported by their cites","value":"77 / 80","unit":null,"n":80,"split":"test","note":"3 partly, 0 not supported, 0 misleading; Claude Code Opus 5.5, blind, saw at most six cites per line"},{"name":"Lines held back for the lawyer","value":"13 / 135","unit":null,"n":135,"split":"test","note":"4 right to hold, 3 dates outside what the line cites, 6 over-strict (3 from a pronoun issue since fixed; 0 over-strict in the post-fix split)"}],"dataset":"6 held-out and 3 post-fix synthetic matters from the chronology generator (records, itemized bills with planted problems, a signed client statement), plus 2 matters written blind by another agent (38 pages); each run end to end: chronology, the client's recorded consent, then the draft.","held_out":true,"caveats":["Synthetic matters only; the generated ones share an author and a template family with the tool.","The blind matters were written by another agent but rendered with the same page renderer.","Answer keys are page-level; box placement comes from the chronology, whose own eval measured it.","The lawyer's review time is an estimate, not timed with a lawyer.","Two fixes were made after the first held-out run (the judge's statement wording, re-reads of unsure bill regions); the post-fix split and the blind split were run once after them."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/settlement-video"},"quality_evidence":[{"tier":"lite","label":"Lite · captions or your own recording, no voice models","evidence":[{"metric":"Script, checks and tie-out","value":"same as standard (the voice does not change them)","source":"decosa-api docs/evals/settlement-video.md"},{"metric":"Render time, captions only","value":"not measured separately; frames and encoding took 41-43 s of the 60-68 s renders","source":"decosa-api docs/evals/settlement-video.md, render table"}]},{"tier":"standard","label":"Standard · reader, model, house voices and the consented clone (hosted demo)","evidence":[{"metric":"Cites on the right page, lines that render (6 held-out synthetic matters)","value":"376 / 376","source":"decosa-api docs/evals/settlement-video.md, held-out seeds, 29 Sep 2026"},{"metric":"Cites on the right page, 2 matters written blind by another agent","value":"143 / 144 on review (139 / 144 by the writer's key)","source":"decosa-api docs/evals/settlement-video.md, blind split"},{"metric":"Dates in rendered lines that match the answer key (all splits)","value":"100 / 100","source":"decosa-api docs/evals/settlement-video.md: 53 held out, 17 blind, 30 post-fix"},{"metric":"Every itemized charge read and summed to the cent","value":"11 / 11 matters","source":"decosa-api docs/evals/settlement-video.md: the default total also leaves out charges with no matching visit, which equals the key in 7 / 11 (the chronology missed 1-3 visits in the others)"},{"metric":"Planted bill problems flagged (before the injury, duplicate, no matching visit, statement total off)","value":"11/11, 11/11, 11/11, 10/11","source":"decosa-api docs/evals/settlement-video.md, all splits"},{"metric":"Rendered lines a blind judge found supported by their cited text","value":"77 / 80 (3 partly, 0 not supported)","source":"decosa-api docs/evals/settlement-video.md, Claude Code Opus 5.5 blind, held-out sample"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":103504,"p95_ms":110047,"runs":5,"receipts_per_run":23,"cost_per_run_usd":0.00961},"selfhost":{"date":"2026-09-29","result":"pass","method":"fresh clone of the branch into a clean directory, run with the host's Python environment (no container build: the server's root disk was full at the time), direct route to the local Qwen3.8-27B, the running document reader, local signing, real-matters mode on","notes":"The rehearsal bundle passed 11/11 in 38.5 s (draft, tie-out flags, render, record verified and failed when changed) and the smoke module passed in 19.9 s."},"known_limits":["Hosted numbers are the whole sample task measured on production (draft + one line re-checked + render with a house voice, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","Measured on synthetic matters only; real records, bills and photos are not measured, and the lawyer's review time (10-20 minutes estimated) has not been timed with a lawyer.","When the chronology misses a visit, a real charge on that day is flagged as having no matching visit and left out of the default total until the lawyer puts it back (4 of 11 eval matters, $186-594).","A printed statement total on a degraded scan or fax is sometimes unreadable; the tool says so and the video uses the sum of the lines.","A child's photo is refused until a guardian consent path for this purpose is approved.","No causation, future-care, wage-loss or billed-versus-paid figures: only what the chronology and the itemized bills hold."],"receipt_coverage":"full"},"cost_per_run_usd":0.00961,"rehearsal_bundle":{"url":"/samples/settlement-video.zip","checks":11,"bytes":1951},"models":[{"name":"decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks","role":"The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure of","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Docling 2.130 with the Heron layout model (document reader block)","role":"Finds the regions of each bill and statement page (tables, text) with their boxes","license":"MIT (Docling) + Apache-2.0 (weights)","hf_repo":"docling-project/docling-layout-heron"},{"name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","role":"Reads the bill tables as cells and the statement's paragraphs","license":"Apache-2.0","hf_repo":"PaddlePaddle/PaddleOCR-VL-1.6"},{"name":"Kokoro-82M (stock voicepacks am_michael, af_heart)","role":"The stock house voice that reads the approved script, when the lawyer picks it","license":"Apache-2.0","hf_repo":"hexgrad/Kokoro-82M"},{"name":"Chatterbox Multilingual","role":"The attorney's own voice, generated from the attorney's consent recording, only with an active consent-ledger entry for this matter","license":"MIT","hf_repo":"ResembleAI/chatterbox"}],"licence":"permissive","links":{"metrics":"/metrics/settlement-video","page":"/legal/settlement-video","json":"/use-cases/settlement-video.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"interview-themes","num":"174","name":"Interview themes","status":"preview","industries":["science-research"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Words credited to the wrong speaker (12 public-domain interviews)","value":"0.05% (22 of 41,624)","unit":null,"n":41624,"split":"test","note":"Interviewer words put in a participant's mouth: 4. Without linking voices across 10-minute pieces: 4.0% (1,269 interviewer words credited to participants; worst interview 31%)."},{"name":"Agreement with published human coding (Cohen's κ, 28 codes, test split)","value":"0.49","unit":null,"n":null,"split":"test","note":"All 1,000 passages: 0.52. With code names only as definitions: 0.29."},{"name":"Agreement with published human coding (Cohen's κ, 28 codes, all passages)","value":"0.52","unit":null,"n":1000,"split":"dev","note":null},{"name":"Our coder vs a blind second coder (Cohen's κ)","value":"0.71","unit":null,"n":120,"split":"dev","note":"The blind second coder is a model playing a careful researcher, coding by hand."},{"name":"Blind second coder vs the published coding, for comparison (Cohen's κ)","value":"0.62","unit":null,"n":120,"split":"dev","note":null},{"name":"Three coders: published, blind, ours (Fleiss' κ)","value":"0.61","unit":null,"n":120,"split":"dev","note":null},{"name":"Quotes passing the word-for-word and speaker check","value":"15 of 15","unit":null,"n":15,"split":"test","note":"Recorded sample run of 12 interviews. Planted test on 100 real quotes: a changed word, a dropped word, interviewer words and a wrong speaker were each caught 100 of 100."}],"dataset":"Speaker attribution: 12 episodes of NASA's Houston We Have a Podcast (US government work, public domain), first 21 minutes each (4.2 h), scored word by word against NASA's published transcripts. Coding: Knowledge Exchange PRRO interview coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), 1,000 human-coded passages against a 28-code hierarchy, shuffled; a 120-passage sample coded blind by a second coder.","held_out":false,"caveats":["Attribution was measured on clean studio recordings with one guest; a simulated video-call copy of 4 interviews scored 0.03%, but real noisy calls, crosstalk and similar voices were not tested.","Two speaker-linking rules were designed after seeing extra speaker ids on the test interviews (thresholds were set on 2 dev episodes).","Coding agreement is one published dataset in one field; the definitions variant was chosen after the names-only run on the same passages, so the kappa is not held out.","The blind second coder is a model playing a careful researcher, not a person.","Agreement drops to 0.29 when codes have names only: definitions matter.","Theme quality against a human thematic analysis was not measured."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/interview-themes"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"agreement with human coding","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"Words credited to the wrong speaker, 12 public-domain interviews (41,624 words)","value":"0.05% (4.0% without voice linking)","source":"decosa-api docs/evals/interview-themes.md, measured 2026-09-29"},{"metric":"Cohen's kappa with published human coding, 28 codes, 1,000 passages (test split)","value":"0.52 (0.49)","source":"decosa-api interview-themes eval, PRRO coding (Zenodo 10.5281/zenodo.5512420, CC BY 4.0), measured 2026-09-29"},{"metric":"a blind second coder vs the published coding, 120 passages (for comparison)","value":"0.62","source":"same eval, 2026-09-29"}]},{"tier":"wanted","label":"Wanted · a much larger coder on your own hardware","evidence":[{"metric":"agreement with human coding","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":69449,"p95_ms":75043,"runs":5,"receipts_per_run":62,"cost_per_run_usd":0.098598},"selfhost":{"date":"2026-09-29","result":"partial","method":"the branch's API run directly on a GPU server with DECOSA_THEMES_SAMPLES_ONLY=0 against the local model servers (not a fresh compose)","notes":"Pasted transcripts, codebook, coding, themes, own model and every export worked, and the rehearsal bundle passed 8 of 8; the docker compose in the assemble prompt was not run end to end."},"known_limits":["Hosted numbers are the whole sample task measured on production (codebook proposal, approval, then coding and themes for all 12 sample interviews, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.","Transcription runs about 14x faster than real time on one shared GPU: about 4 to 5 minutes per interview hour, one recording at a time.","The hosted demo takes the public-domain sample interviews only."],"receipt_coverage":"full"},"cost_per_run_usd":0.098598,"rehearsal_bundle":{"url":"/samples/interview-themes.zip","checks":8,"bytes":32632},"models":[{"name":"MOSS-Transcribe-Diarize 0.9B","role":"Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute piece","license":"Apache-2.0","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize"},{"name":"ECAPA-TDNN speaker embeddings (ONNX export)","role":"Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)","license":"Apache-2.0","hf_repo":"speechbrain/spkrec-ecapa-voxceleb"},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are code","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"bge-small-en-v1.5 (ONNX)","role":"Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)","license":"MIT","hf_repo":"BAAI/bge-small-en-v1.5"}],"licence":"permissive","links":{"metrics":"/metrics/interview-themes","page":"/tools/research/interview-themes","json":"/use-cases/interview-themes.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"mix-cue-sheet","num":"180","name":"Mix cue sheet","status":"live","industries":["music","creative-media"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Song starts within 10 s, with your track files","value":"105/105 (92/105 within 5 s, 24/105 within 1 s)","unit":null,"n":105,"split":"test","note":"105 starts predicted. Time per 5-minute mix: p50 4.46 s, p95 6.26 s."},{"name":"Song starts within 10 s, titles only (the hosted default)","value":"84/105 (51/105 within 5 s, 15/105 within 1 s)","unit":null,"n":105,"split":"test","note":"The tracklist without times. Time per 5-minute mix: p50 0.85 s, p95 1.25 s."},{"name":"Song starts within 1 / 5 / 10 s, audio only","value":"3/105 / 10/105 / 19/105","unit":null,"n":105,"split":"test","note":"Only 20 starts predicted: not useful on its own."},{"name":"By join, with your track files (within 5 s / 10 s, of 35 each)","value":"cuts 23 / 31, crossfades 31 / 35, blends 33 / 35","unit":null,"n":105,"split":"test","note":null},{"name":"By join, titles only (within 5 s / 10 s, of 35 each)","value":"cuts 26 / 30, crossfades 11 / 30, blends 14 / 24","unit":null,"n":105,"split":"test","note":null},{"name":"Learned detector on (self-host): within 10 s, titles only / with track files","value":"61/105 / 85/105","unit":null,"n":105,"split":"test","note":"Not better on these short synthetic songs; it helps on real long mixes (dev set)."},{"name":"Starts within 10 / 20 / 30 s, titles only, detector off (hosted setting)","value":"61 / 75 / 86 of 190","unit":null,"n":190,"split":"dev","note":"Within 5 s: 40. No track files and no times: every start is placed from the audio."},{"name":"Starts within 10 / 20 / 30 s, titles only, detector on","value":"57 / 73 / 81 of 190","unit":null,"n":190,"split":"dev","note":null},{"name":"Starts within 10 / 20 / 30 s, audio only, detector off / on","value":"59 / 70 / 81 vs 78 / 104 / 129 of 190","unit":null,"n":190,"split":"dev","note":"With the detector on, 358 starts were predicted for 190 real ones."},{"name":"Starts within 10 / 20 / 30 s with a fingerprint service's identifications, titles and the detector","value":"99 / 123 / 144 of 190","unit":null,"n":190,"split":"dev","note":"Within 1 s: 26; within 5 s: 74; 277 predicted starts. Detector off: 70 / 95 / 146. The identifications came from a fingerprint service that is not part of the product. This setting reproduces the original engine's numbers exactly."}],"dataset":"Held out: 20 synthetic mixes built from 19 of Decosa's own Make a song songs (60 s each, similarity check clear), 105 song starts, 99 minutes of audio; each song played at 0.96-1.04 speed and trimmed 0-6 s; joins are cuts, crossfades (4-12 s) and bass-swap blends (12-24 s), 35 of each. Development: 8 public DJ mixes with 190 song starts from their published timed tracklists (mixes and DJs not named).","held_out":true,"caveats":["The held-out mixes are synthetic, built from 60-second songs, not real DJ mixes.","The same author wrote the mixer and the engine.","Each test song appears in about 5 of the 20 mixes.","Every threshold was frozen on a separate dev split before the test set was built; the test set was run once.","The public-mix numbers are a development set, not held out: the original engine's settings were chosen while looking at them. Real DJ mixes with long blends are harder (61 of 190 within 10 s from titles alone).","DJs' own timestamps disagree by about 9 s, so 10 s is the practical bar.","The runs with identifications used cached results from a fingerprint service that is not part of the product."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/mix-cue-sheet"},"quality_evidence":[{"tier":"lite","label":"Lite · titles and times only, any CPU","evidence":[{"metric":"Song starts within 1 / 5 / 10 s, titles only (held-out synthetic mixes)","value":"15/105 / 51/105 / 84/105","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (held out: 20 synthetic mixes of Decosa's own songs, 105 song starts, run once with thresholds frozen on a dev split)"},{"metric":"Song starts within 10 / 20 / 30 s, titles only, detector off, 8 real DJ mixes","value":"61 / 75 / 86 of 190 (5 s: 40)","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (8 public DJ mixes, 190 song starts; development numbers, not held out)"}]},{"tier":"standard","label":"Standard · adds your own track files (hosted)","evidence":[{"metric":"Song starts within 1 / 5 / 10 s, with your track files (held-out synthetic mixes)","value":"24/105 / 92/105 / 105/105, with 105 starts predicted","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (held out: 20 synthetic mixes of Decosa's own songs, 105 song starts, run once with thresholds frozen on a dev split)"},{"metric":"By join (within 5 s / 10 s, of 35 each), with your track files","value":"cuts 23 / 31, crossfades 31 / 35, blends 33 / 35","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (held out: 20 synthetic mixes of Decosa's own songs, 105 song starts, run once with thresholds frozen on a dev split)"},{"metric":"Time per 5-minute mix with track files (engine, held-out set)","value":"p50 4.46 s, p95 6.26 s","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (held out: 20 synthetic mixes of Decosa's own songs, 105 song starts, run once with thresholds frozen on a dev split)"}]},{"tier":"best","label":"Best · adds the learned transition detector (self-host only)","evidence":[{"metric":"Held-out synthetic mixes, detector on: song starts within 10 s, titles only / with track files","value":"61/105 / 85/105 (not better on short synthetic songs)","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (held out: 20 synthetic mixes of Decosa's own songs, 105 song starts, run once with thresholds frozen on a dev split)"},{"metric":"Titles, the detector and a fingerprint service's identifications, 8 real DJ mixes (helps on real long mixes)","value":"99 / 123 / 144 of 190 within 10 / 20 / 30 s (1 s: 26, 5 s: 74; 277 predicted starts); detector off: 70 / 95 / 146","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (8 public DJ mixes, 190 song starts; development numbers, not held out)"},{"metric":"Audio only (no titles), detector on, 8 real DJ mixes","value":"78 / 104 / 129 of 190 within 10 / 20 / 30 s (358 predicted starts)","source":"decosa-api docs/evals/mix-cue-sheet.md, 2026-09-29 (8 public DJ mixes, 190 song starts; development numbers, not held out)"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":6960,"p95_ms":7340,"runs":5,"receipts_per_run":0,"cost_per_run_usd":0},"selfhost":null,"known_limits":["It names only songs you give a title or a file for; there is no catalogue lookup on the hosted tool.","Titles without times or files are placed from the audio alone and marked estimated: 84 of 105 within 10 s on held-out synthetic mixes, but 61 of 190 on real DJ mixes with long blends.","Audio alone, with no titles and no files, finds few starts (19 of 105 within 10 s); it is not offered as a mode on its own.","The learned transition detector is off on the hosted tool (unlicensed training audio)."],"receipt_coverage":"none"},"cost_per_run_usd":0,"rehearsal_bundle":{"url":"/samples/mix-cue-sheet.zip","checks":12,"bytes":1330},"models":[{"name":"decosa-cue engine (decosa_api/cue)","role":"Song starts: spectral features and a novelty curve from the audio, a tracklist parser, and a resolver that turns your times, your track-file matches and the audio's change points into one start per song; writes the chapters (no model; CPU)","license":"Apache-2.0","hf_repo":null},{"name":"Landmark matcher for your own track files (decosa_api/verticals/clearance fingerprinter)","role":"Finds each of your own track files in the mix from the audio (landmark pairs, searched across tempo changes of up to 6% and the pitch shift that comes with them) so its song gets an exact start (no model; CPU)","license":"AGPL-3.0-or-later","hf_repo":null}],"licence":"permissive","links":{"metrics":"/metrics/mix-cue-sheet","page":"/tools/media/mix-cue-sheet","json":"/use-cases/mix-cue-sheet.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"review-reply","num":"182","name":"Review reply with patient privacy","status":"live","industries":["healthcare","sales-marketing"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Healthcare replies that confirm a patient (blind judge)","value":"0 / 40","unit":null,"n":40,"split":"test","note":"held-out set generated after the guard was frozen; target 0"},{"name":"Replies a business could post as written (blind judge)","value":"93 / 100","unit":null,"n":100,"split":"test","note":"healthcare 38/40, other 55/60"},{"name":"Healthcare replies that fell back to the fixed safe reply","value":"24 / 40","unit":null,"n":40,"split":"test","note":"the price of blocking broadly"},{"name":"Replies with an offer or a fact nobody gave (blind judge)","value":"4 / 100","unit":null,"n":100,"split":"test","note":null},{"name":"First guard on the development set: healthcare replies that confirm a patient","value":"9 / 40","unit":null,"n":40,"split":"dev","note":"code rules only; led to the model privacy check and broader rules"}],"dataset":"200 synthetic reviews (80 for health and care businesses) written by Qwen3.8-27B from seeded plans: set 1 for development, set 2 held out and generated after the guard was frozen. Replies judged blind by Claude Code (Opus), which saw only the business type, stars, review and reply.","held_out":true,"caveats":["One judge (a frontier model), no human labels.","Synthetic reviews written by the same model family that drafts the replies.","English only.","The guard blocks broadly, so many health and care replies are the generic fixed reply."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/review-reply"},"quality_evidence":[{"tier":"lite","label":"Lite · one 48 GB card","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"standard","label":"Standard · the hosted demo, one 96 GB card","evidence":[{"metric":"held-out healthcare replies that confirm a patient (blind judge)","value":"0 of 40","source":"decosa-api docs/evals/review-reply.md, gateway route, 29 Sep 2026"},{"metric":"held-out replies a business could post as written (blind judge)","value":"93 of 100","source":"decosa-api docs/evals/review-reply.md"}]},{"tier":"best","label":"Best · DeepSeek-V4-Flash on two more cards","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]},{"tier":"wanted","label":"Wanted · two large judges from different families","evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":1801,"p95_ms":4237,"runs":100,"receipts_per_run":1.55,"cost_per_run_usd":0.000441},"selfhost":null,"known_limits":["Timings and cost were measured on the pre-release server through the production gateway. On production (30 Sep 2026) the samples, a made-up review typed in by hand and the own-reply check were run end to end in a browser; self-hosting from the assemble prompt has not been verified yet.","Reviews are synthetic, written by the same model family that drafts the replies; real reviews are messier.","English only.","A new phrasing that confirms a patient can still slip past both checks; read every reply before posting."],"receipt_coverage":"full"},"cost_per_run_usd":0.000441,"rehearsal_bundle":{"url":"/samples/review-reply.zip","checks":12,"bytes":1499},"models":[{"name":"@decosa/site-kit review guard","role":"The review-reply guard: offers, contact details and numbers not given, placeholders, arguing, and the patient-privacy rules. Deterministic code, the same in TypeScript and Python.","license":"Apache-2.0","hf_repo":null},{"name":"Qwen3.8-27B (NVIDIA NVFP4)","role":"Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"}],"licence":"permissive","links":{"metrics":"/metrics/review-reply","page":"/tools/operations/review-reply","json":"/use-cases/review-reply.json","nightly":"https://api.decosa.ai/verify/status"}},{"id":"what-studies-found","num":"183","name":"What studies found","status":"preview","industries":["science-research","healthcare"],"deploy":["hosted","selfhost"],"eval_summary":{"metrics":[{"name":"Verdict matches the systematic review's (benefit, harm or no claim)","value":"231 / 289 (80%)","unit":null,"n":289,"split":"heldout","note":"95% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289."},{"name":"Calls a benefit where the review found no clear difference","value":"1 / 61 (2%)","unit":null,"n":61,"split":"heldout","note":"The engine it replaced, on the same pairs: 8 / 61."},{"name":"Harms surfaced: the verdict is Worsens, or the table carries its harm flag","value":"36 / 40 (90%)","unit":null,"n":40,"split":"heldout","note":"By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40."},{"name":"Harm verdict or harm flag where the review reports no harm (the price of the flag)","value":"46 / 249 (18%)","unit":null,"n":249,"split":"heldout","note":"A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging."},{"name":"Finds the benefit where the review found one","value":"56 / 82 (68%)","unit":null,"n":82,"split":"heldout","note":"The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear."},{"name":"Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear)","value":"180 / 282 (64%)","unit":null,"n":282,"split":"heldout","note":"Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282."},{"name":"The review behind the label was among the studies read","value":"256 / 290 (88%)","unit":null,"n":290,"split":"heldout","note":"Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290."},{"name":"One abstract read alone: verdict matches the label (benefit, harm or no claim)","value":"290 / 331 (88%)","unit":null,"n":331,"split":"heldout","note":"Harms surfaced per study: 60 / 62."}],"dataset":"290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.","held_out":true,"caveats":["The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.","Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.","Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.","The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.","Abstracts only; no risk-of-bias assessment.","Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts."],"date":"2026-09-30","doc_url":"https://decosa.ai/metrics/evals/what-studies-found"},"quality_evidence":[{"tier":"lite","label":"Lite · one 32 GB card, no reranker","evidence":[{"metric":"Table quality without the reranker","value":"not measured yet","source":"no run without the reranker has been scored"}]},{"tier":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","evidence":[{"metric":"Held-out: the table's verdict matches the systematic review's (benefit, harm or no claim)","value":"231 / 289 (80%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: calls a benefit where the review found no clear difference","value":"1 / 61 (2%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: harms surfaced by the verdict or the harm flag","value":"36 / 40 (90%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: harm verdict or flag where the review reports no harm","value":"46 / 249 (18%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: finds the benefit where the review found one","value":"56 / 82 (68%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"}]}],"benchmark":null,"verification":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":26328,"p95_ms":31027,"runs":5,"receipts_per_run":32,"cost_per_run_usd":0.022935},"selfhost":null,"known_limits":["Hosted: measured on production on 30 Sep 2026 with the first sample (probiotics and antibiotic-associated diarrhea) through the API, one run at a time. Well-studied pairs like the samples read more reviews and cost more than the average table.","Reads abstracts only; results reported only in the full paper are missed.","The one-line verdict is not a systematic review. On held-out pairs it matches the published review's more often than not but not always (the figures are in the eval); the misses are mostly tables that come out mixed or unclear where the review found a benefit.","The rubric puts the newest Cochrane review first even when it studied a narrower group than the question (the melatonin sample: a review of shift workers decides, the other reviews disagree, and the verdict is \"Mixed results\"). Choosing the review by who it studied is planned, not built.","The same search can give a different verdict on another run: the model's calls on what is relevant vary a little, and PubMed changes.","The harm flag errs toward flagging: on held-out pairs it also appeared on tables whose review reports no harm (the figure is in the eval). It is a prompt to read the study.","The band has no risk-of-bias assessment; it is our rubric over the abstracts found, not a GRADE rating.","A meta-analysis and trials it already pooled can sit in the same table."],"receipt_coverage":"full"},"cost_per_run_usd":0.022935,"rehearsal_bundle":{"url":"/samples/what-studies-found.zip","checks":14,"bytes":1457},"models":[{"name":"decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)","role":"Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)","license":"Apache-2.0 (engine); AGPL-3.0-or-later (routes)","hf_repo":null},{"name":"Qwen3.8-27B (NVFP4)","role":"Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstract","license":"Apache-2.0","hf_repo":"nvidia/Qwen3.8-27B-NVFP4"},{"name":"Qwen3-Reranker-4B (evidence retrieval block)","role":"Ranks the PubMed results by relevance to the supplement and outcome before any abstract is read","license":"Apache-2.0","hf_repo":"Qwen/Qwen3-Reranker-4B"}],"licence":"permissive","links":{"metrics":"/metrics/what-studies-found","page":"/tools/life-sciences/what-studies-found","json":"/use-cases/what-studies-found.json","nightly":"https://api.decosa.ai/verify/status"}}]}