17 · Any industry · Compliance and trust · live
Grounding check
Eval results
Scored on a held-out or test splitRun 24 Sep 2026Eval write-up (decosa-api, access required)
- Unsupported-sentence precision / recall, default gate0.593 / 0.556held outn = 2,069F1 0.574, agreement 0.929, Cohen's kappa 0.535, flag rate 8.1%.
- Recall, strict gate (partial counts too)0.966held outn = 2,069Precision 0.276, flag rate 30.1%.
- Response-level F1, default gate0.691held outn = 300P 0.706, R 0.675.
- Baseline: NLI cross-encoder on CPU (lite tier), F10.227held outn = 2,069P 0.133, R 0.792, flag rate 51.4%.
- Default gate precision / recall on dev (tuned on)0.488 / 0.494dev (tuned on)n = 927
Dataset
RAGTruth (MIT): human span labels on responses from six LLMs to QA, news summaries and data-to-text; dev 120 responses from train, test 300 responses (2,069 sentences, 178 unsupported) from the test split, run once after freezing.
Caveats
- One benchmark, English only, generated by 2023-era models; the false-flag rate must be measured per domain before the gate blocks without a human look.
- A date line was added to the prompt after the dev run and before the test run.
- The eval ran on the direct route to the same model server, so its calls carry no gateway receipts.
- RAGTruth annotators are lenient on added detail, so some strict-gate flags count against the checker without being wrong; the labels were not re-annotated.
- 'Supported by the sources' is not 'true': a sentence copied from a wrong source passes.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 17 s
- Receipts
- 5
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 25 Sep 2026: Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers
Verified on 2026-09-25, option A: the api starts, the thermostat sample blocks with the battery and 240 V sentences contradicted and the reset sentence supported at S1.3, the stream and the signed report behave as documented (p50 1.9 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Option B (CPU NLI judge) also ran: it blocked the sample but marked the battery-life sentence supported.
Rehearsal bundle: grounding.zip (2 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- The judge reads only the sources given: "supported" is not "true".
- Each sentence is one model call that re-sends the sources, so tokens grow with sentences x source length (about 5.8k for a five-sentence answer against a one-page manual).
- The CPU NLI judge (self-host option B) is coarser: on the thermostat sample it missed one of the two contradictions.
- Sources must be pasted text; a bare URL is refused (nothing is fetched).
- Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Checker: segmentation, evidence selection, gate, signed report (no model; CPU)decosa-api grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
- Judge: one call per sentence, verdict and cited spansQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · CPU only, no GPU (3)
- RAGTruth test, sentence level: precision / recall on unsupported: 0.133 / 0.792 (F1 0.227)docs/evals/grounding.md, 300 held-out responses, threshold from dev
- RAGTruth test: agreement with human labels: 53.6% (κ 0.09)docs/evals/grounding.md
- Lexical-overlap baseline, same test: 0.177 / 0.455 (F1 0.255), agreement 77.1%docs/evals/grounding.md
Standard · one GPU for the judge (hosted demo) (4)
- RAGTruth test, default gate (block = unsupported or contradicted): precision / recall on unsupported: 0.593 / 0.556 (F1 0.574)docs/evals/grounding.md, 2,069 sentences, 178 unsupported, held out
- RAGTruth test, default gate: agreement with human labels: 92.9% (κ 0.535)docs/evals/grounding.md
- RAGTruth test, strict (partial also flagged): precision / recall: 0.276 / 0.966 (F1 0.429), agreement 77.9%docs/evals/grounding.md
- RAGTruth test, response level, default gate: precision / recall: 0.706 / 0.675docs/evals/grounding.md
Wanted · a panel of the largest open judges (1)
- sentence-level F1 on the grounding set, same protocol as standard: not measured yet