Skip to content
decosa

17 · Any industry · Compliance and trust · live

Grounding check

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 24 Sep 2026Eval write-up (decosa-api, access required)

  • Unsupported-sentence precision / recall, default gate0.593 / 0.556held outn = 2,069F1 0.574, agreement 0.929, Cohen's kappa 0.535, flag rate 8.1%.
  • Recall, strict gate (partial counts too)0.966held outn = 2,069Precision 0.276, flag rate 30.1%.
  • Response-level F1, default gate0.691held outn = 300P 0.706, R 0.675.
  • Baseline: NLI cross-encoder on CPU (lite tier), F10.227held outn = 2,069P 0.133, R 0.792, flag rate 51.4%.
  • Default gate precision / recall on dev (tuned on)0.488 / 0.494dev (tuned on)n = 927

Dataset

RAGTruth (MIT): human span labels on responses from six LLMs to QA, news summaries and data-to-text; dev 120 responses from train, test 300 responses (2,069 sentences, 178 unsupported) from the test split, run once after freezing.

Caveats

  • One benchmark, English only, generated by 2023-era models; the false-flag rate must be measured per domain before the gate blocks without a human look.
  • A date line was added to the prompt after the dev run and before the test run.
  • The eval ran on the direct route to the same model server, so its calls carry no gateway receipts.
  • RAGTruth annotators are lenient on added detail, so some strict-gate flags count against the checker without being wrong; the labels were not re-annotated.
  • 'Supported by the sources' is not 'true': a sentence copied from a wrong source passes.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
17 s
Receipts
5
Model calls
n/a
Tokens
n/a
Cost per run
$0.002

Self-host verification

Verified on 25 Sep 2026: Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers

Verified on 2026-09-25, option A: the api starts, the thermostat sample blocks with the battery and 240 V sentences contradicted and the reset sentence supported at S1.3, the stream and the signed report behave as documented (p50 1.9 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Option B (CPU NLI judge) also ran: it blocked the sample but marked the battery-life sentence supported.

Rehearsal bundle: grounding.zip (2 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • The judge reads only the sources given: "supported" is not "true".
  • Each sentence is one model call that re-sends the sources, so tokens grow with sentences x source length (about 5.8k for a five-sentence answer against a one-page manual).
  • The CPU NLI judge (self-host option B) is coarser: on the thermostat sample it missed one of the two contradictions.
  • Sources must be pasted text; a bare URL is refused (nothing is fetched).
  • Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Checker: segmentation, evidence selection, gate, signed report (no model; CPU)decosa-api grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
  • Judge: one call per sentence, verdict and cited spansQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · CPU only, no GPU (3)
  • RAGTruth test, sentence level: precision / recall on unsupported: 0.133 / 0.792 (F1 0.227)docs/evals/grounding.md, 300 held-out responses, threshold from dev
  • RAGTruth test: agreement with human labels: 53.6% (κ 0.09)docs/evals/grounding.md
  • Lexical-overlap baseline, same test: 0.177 / 0.455 (F1 0.255), agreement 77.1%docs/evals/grounding.md
Standard · one GPU for the judge (hosted demo) (4)
  • RAGTruth test, default gate (block = unsupported or contradicted): precision / recall on unsupported: 0.593 / 0.556 (F1 0.574)docs/evals/grounding.md, 2,069 sentences, 178 unsupported, held out
  • RAGTruth test, default gate: agreement with human labels: 92.9% (κ 0.535)docs/evals/grounding.md
  • RAGTruth test, strict (partial also flagged): precision / recall: 0.276 / 0.966 (F1 0.429), agreement 77.9%docs/evals/grounding.md
  • RAGTruth test, response level, default gate: precision / recall: 0.706 / 0.675docs/evals/grounding.md
Wanted · a panel of the largest open judges (1)
  • sentence-level F1 on the grounding set, same protocol as standard: not measured yet

How we measure · All tools