Grounding check: eval on RAGTruth (24 Sep 2026)
Question: does the checker flag the sentences that human annotators marked as not supported by the sources?
Data
- RAGTruth (Niu et al., 2024; github.com/ParticleMedia/RAGTruth, MIT licence). Responses from six LLMs to QA (MS MARCO passages), news summarisation (CNN/DM) and data-to-text (Yelp business JSON), with human span labels: Evident/Subtle Conflict and Evident/Subtle Baseless Info.
- Dev: 40 responses per task sampled from the train split (seed 17): 120 responses, 927 sentences, 83 unsupported. Used for prompt development, the baseline thresholds and the confidence table.
- Test (held out): 100 responses per task from the test split (seed 17): 300 responses, 2,069 sentences, 178 unsupported (8.6 %). Run once, after the prompt and thresholds were frozen.
- Sentence labels: the checker's own splitter (
segment.py) splits each response; a sentence is unsupported when any annotated span overlaps it. "Flagged" means any verdict other than supported or no_claim. - Sources as sent to the checker: QA passages as separate sources plus the question; the article for summaries; the business JSON pretty-printed (one field per line) for data-to-text.
Results (test, sentence level)
| System | Precision (unsupported) | Recall | F1 | Agreement | Cohen's κ | Flag rate |
|---|---|---|---|---|---|---|
| Qwen3.8-27B judge, default gate (block = unsupported or contradicted) | 0.593 | 0.556 | 0.574 | 0.929 | 0.535 | 8.1 % |
| Qwen3.8-27B judge, strict (partial counts too) | 0.276 | 0.966 | 0.429 | 0.779 | 0.341 | 30.1 % |
| Judge + NLI cross-check (flag only if the NLI model does not find it entailed, threshold 0.68 from dev) | 0.304 | 0.882 | 0.452 | 0.816 | 0.371 | 25.0 % |
Baseline: NLI cross-encoder on CPU (cross-encoder/nli-deberta-v3-base, best 4 BM25 windows, threshold from dev) |
0.133 | 0.792 | 0.227 | 0.536 | 0.094 | 51.4 % |
| Baseline: lexical overlap (content words and every number, threshold 0.36 from dev) | 0.177 | 0.455 | 0.255 | 0.771 | 0.149 | 22.1 % |
| Flag everything | 0.086 | 1.000 | 0.158 | 0.086 | 0 | 100 % |
Per task, default gate: data-to-text P 0.762 / R 0.582; QA P 0.385 / R 0.541; summaries P 0.484 / R 0.484. Response level (a response is flagged when any sentence is): default gate P 0.706, R 0.675, F1 0.691; strict P 0.489, R 0.991; NLI baseline P 0.436, R 0.982. When the judge says contradicted, the annotators marked a Conflict span on that sentence 58 % of the time (115 predicted).
Dev (for reference, same code): default gate P 0.488 / R 0.494; strict P 0.279 / R 0.928.
Reading the numbers
- Two operating points, one model. Counting
partialfinds almost every annotated sentence (97 %) but flags 30 % of sentences. Most of those extra flags are added details the annotators left alone: "making it an ideal place for gatherings", "a great spot for hearty meals", procedural filler in how-to answers. Some are real misses in the labels, but we did not re-annotate, so they count against us. The default gate therefore blocks onlyunsupportedandcontradictedand only flagspartial. - The judge beats both CPU baselines by a wide margin at every operating point. The NLI cross-encoder is the lite tier for machines without a GPU; it is much weaker, and the numbers say so.
- Confidence is per verdict, measured on dev: supported 99 %, unsupported 60 %, contradicted 42 %, partial 20 %, no_claim 98 % (the judge nearly always states "high" certainty, so certainty adds little). This is agreement with the RAGTruth annotators, who are lenient on added detail, not a probability that the sentence is false.
Honest caveats
- Tuned on dev, not on test. Two prompt revisions were tried on dev (the first counted "the passages do not answer this" as a claim and over-called PARTIAL; the second narrowed PARTIAL to specific facts and unfounded praise). A third change added today's date to the prompt, for durations such as "since 2021", after the dev run and before the test run; the test run uses the shipped prompt. The confidence table comes from the dev run.
- Route. The eval ran on the direct route to the same vLLM server the gateway uses (
qwen3.8-27b, temperature 0), because the shared gateway was saturated by other evals at the time. Model, prompt and parameters are the ones the hosted API uses; the eval calls have no gateway receipts (except the first 20 dev responses, which ran through the gateway before the switch). The fixture runs on the site went through the gateway. - One benchmark, English only, generated by 2023-era models. The false-flag rate must be measured per domain before the gate blocks without a human look (page 26 already warned about this for the clinical verifier).
- "Supported by the sources" is not "true". A sentence copied from a wrong source passes.
Reproduce
python scripts/grounding_eval.py build --data <RAGTruth/dataset> --work ~/.cache/grounding-eval # 120 dev + 300 test responses
DECOSA_LLM_ROUTE=direct python scripts/grounding_eval.py llm --work ~/.cache/grounding-eval --split dev
DECOSA_LLM_ROUTE=direct python scripts/grounding_eval.py llm --work ~/.cache/grounding-eval --split test
python scripts/grounding_eval.py nli --work ~/.cache/grounding-eval # needs torch (CPU) and transformers
python scripts/grounding_eval.py score --work ~/.cache/grounding-eval --calibration decosa_api/verticals/grounding/data/calibration.json
Cost on our server: dev 740 judge calls in 202 s, test 1,995 calls in 679 s (6 in flight, GPU1 shared with other work; 3.3 M prompt tokens, 138 k generated). NLI baseline: 2,996 sentences in 658 s on 8 CPU threads.