{"schema_version":"1","site":"https://decosa.ai","id":"grounding","num":"17","name":"Grounding check","tool_name":"Catch AI answers your sources don't back","short":"Grounding","blurb":"Checks every sentence of a draft against the sources you give it: supported, partial, unsupported or contradicted, with the source text it rests on. A claims gate can stop an email or page before it ships.","status":"live","labels":{"industry":["general","compliance-trust"],"job":["review"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["data","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["general","compliance-trust"],"runs_in":["hosted","selfhost","crp-deal-research"],"part_of":[],"built_from":["grounding","signed-record"],"models":"Qwen3.8-27B","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the checker itself runs on CPU","final_artifact":"Per-sentence verdicts with source spans, a gate decision and a signed report.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per verdict; signed report","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":17202,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0021},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers","notes":"Verified on 2026-09-25, option A: the api starts, the thermostat sample blocks with the battery and 240 V sentences contradicted and the reset sentence supported at S1.3, the stream and the signed report behave as documented (p50 1.9 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Option B (CPU NLI judge) also ran: it blocked the sample but marked the battery-life sentence supported."},"known_limits":["The judge reads only the sources given: \"supported\" is not \"true\".","Each sentence is one model call that re-sends the sources, so tokens grow with sentences x source length (about 5.8k for a five-sentence answer against a one-page manual).","The CPU NLI judge (self-host option B) is coarser: on the thermostat sample it missed one of the two contradictions.","Sources must be pasted text; a bare URL is refused (nothing is fetched).","Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Unsupported-sentence precision / recall, default gate","value":"0.593 / 0.556","unit":null,"n":2069,"split":"heldout","note":"F1 0.574, agreement 0.929, Cohen's kappa 0.535, flag rate 8.1%."},{"name":"Recall, strict gate (partial counts too)","value":"0.966","unit":null,"n":2069,"split":"heldout","note":"Precision 0.276, flag rate 30.1%."},{"name":"Response-level F1, default gate","value":"0.691","unit":null,"n":300,"split":"heldout","note":"P 0.706, R 0.675."},{"name":"Baseline: NLI cross-encoder on CPU (lite tier), F1","value":"0.227","unit":null,"n":2069,"split":"heldout","note":"P 0.133, R 0.792, flag rate 51.4%."},{"name":"Default gate precision / recall on dev (tuned on)","value":"0.488 / 0.494","unit":null,"n":927,"split":"dev","note":null}],"dataset":"RAGTruth (MIT): human span labels on responses from six LLMs to QA, news summaries and data-to-text; dev 120 responses from train, test 300 responses (2,069 sentences, 178 unsupported) from the test split, run once after freezing.","held_out":true,"caveats":["One benchmark, English only, generated by 2023-era models; the false-flag rate must be measured per domain before the gate blocks without a human look.","A date line was added to the prompt after the dev run and before the test run.","The eval ran on the direct route to the same model server, so its calls carry no gateway receipts.","RAGTruth annotators are lenient on added detail, so some strict-gate flags count against the checker without being wrong; the labels were not re-annotated.","'Supported by the sources' is not 'true': a sentence copied from a wrong source passes."],"date":"2026-09-24","doc_url":"https://decosa.ai/metrics/evals/grounding"},"stack":{"summary":"Send a text and the sources it is supposed to rest on. The checker splits the text into sentences with exact offsets, numbers the source spans, and asks an open judge model about each sentence in its own call: supported, partial, unsupported or contradicted, with the spans it rests on and a confidence measured on a benchmark. The claims gate turns that into pass, flag or block before an email, FAQ or answer ships, and every check ends with an Ed25519-signed report of hashes and offsets. It is the clinical scribe's sentence verifier, made general, and a module other tools can import.","tagline":"Is each sentence supported by the sources? Per-sentence verdicts with the source span, a claims gate, and a signed report.","deployment":"hosted-or-self-host","regulatory_note":"A grounding check is a quality control, not a fact check: it reads only the sources you send, and a sentence copied from a wrong source passes. Verdicts come from a language model and can be wrong in both directions (see the RAGTruth numbers below). The hosted demo keeps no text: drafts and sources stay in memory for the request, the report holds hashes only, and logs carry counts. For unpublished or confidential drafts, self-host. Where rules require substantiating marketing claims (for example the FTC's advertising substantiation policy in the US), a pass here is a record that a check ran against the sources you chose, not proof that a claim is substantiated. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B; cross-encoder/nli-deberta-v3-base). Checked 24 Sep 2026.","components":[{"id":"checker","role":"Checker: segmentation, evidence selection, gate, signed report (no model; CPU)","name":"decosa-api grounding module (decosa_api/verticals/grounding)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; rule sentence splitter, BM25 over source spans","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","wanted"],"alternative_to":null},{"id":"judge","role":"Judge: one call per sentence, verdict and cited spans","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off, prefix caching","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"nli","role":"Lite judge: NLI cross-encoder on CPU (self-host only)","name":"cross-encoder/nli-deberta-v3-base","hf_repo":"cross-encoder/nli-deberta-v3-base","license":"Apache-2.0","params":"184M","quant":"FP32","vram_gb":0,"memory_gb_estimate":null,"engine":"transformers on CPU (torch CPU build)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"gemma-judge","role":"Fast judge with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0","params":"26B","quant":"BF16 (49 GB)","vram_gb":49,"memory_gb_estimate":null,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null},{"id":"dsv41-wanted","role":"Second judge: is each sentence supported by its sources?","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"glm-wanted","role":"Third judge, from another family","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · CPU only, no GPU","summary":"The checker plus an NLI cross-encoder. Runs anywhere; far less accurate, and no model receipts.","components":["checker","nli"],"hardware":"Any Linux or macOS machine with Python 3.11+","quality_evidence":[{"metric":"RAGTruth test, sentence level: precision / recall on unsupported","value":"0.133 / 0.792 (F1 0.227)","source":"docs/evals/grounding.md, 300 held-out responses, threshold from dev"},{"metric":"RAGTruth test: agreement with human labels","value":"53.6% (κ 0.09)","source":"docs/evals/grounding.md"},{"metric":"Lexical-overlap baseline, same test","value":"0.177 / 0.455 (F1 0.255), agreement 77.1%","source":"docs/evals/grounding.md"}],"latency_note":"measured: a fraction of a second per sentence on a few CPU threads.","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"No model call, so no receipt; the report is still signed by the box.","hosting":null},{"id":"standard","label":"Standard · one GPU for the judge (hosted demo)","summary":"Qwen3.8-27B judges each sentence in its own receipted call. This is what the hosted API runs.","components":["checker","judge"],"hardware":"1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"RAGTruth test, default gate (block = unsupported or contradicted): precision / recall on unsupported","value":"0.593 / 0.556 (F1 0.574)","source":"docs/evals/grounding.md, 2,069 sentences, 178 unsupported, held out"},{"metric":"RAGTruth test, default gate: agreement with human labels","value":"92.9% (κ 0.535)","source":"docs/evals/grounding.md"},{"metric":"RAGTruth test, strict (partial also flagged): precision / recall","value":"0.276 / 0.966 (F1 0.429), agreement 77.9%","source":"docs/evals/grounding.md"},{"metric":"RAGTruth test, response level, default gate: precision / recall","value":"0.706 / 0.675","source":"docs/evals/grounding.md"}],"latency_note":"measured: seconds per short check on a quiet gateway; much longer when the shared gateway is busy.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every verdict is a separate gateway call with a gateway-signed receipt; the signed report lists them all.","hosting":null},{"id":"wanted","label":"Wanted · a panel of the largest open judges","summary":"DeepSeek-V4.1-Flash and GLM-5.3-Flash judge each sentence beside Qwen3.8-27B. A sentence counts as supported only when the panel agrees, and disagreements are shown. Not served yet.","components":["checker","judge","dsv41-wanted","glm-wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.","quality_evidence":[{"metric":"sentence-level F1 on the grounding set, same protocol as standard","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[{"id":"fast-judge","label":"Fast option-token judge","components":["gemma-judge"],"hardware":"1x RTX PRO 6000 96 GB (BF16 weights are 49 GB)","use":"A 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /grounding/info, /grounding/samples; POST /grounding/check (SSE or JSON), /grounding/gate, /grounding/verify. Keeps no text."},{"name":"vLLM (judge)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"RAGTruth","url":"https://github.com/ParticleMedia/RAGTruth","license":"MIT","purpose":"The eval: human span labels of unsupported and conflicting content in QA, summary and data-to-text responses. We used a 300-response sample of its test split, held out."},{"name":"scripts/grounding_eval.py and docs/evals/grounding.md","url":null,"license":"Apache-2.0","purpose":"Rebuilds the dev and test samples, runs the judge and both baselines, and writes the metrics and the confidence table."},{"name":"POST /grounding/verify","url":null,"license":"Apache-2.0","purpose":"Checks a report's signature against this server's key and, if you send them, the text and source hashes. The console also checks the signature in your browser with WebCrypto."}],"hardware":[{"tier":"Any CPU, no GPU","fits":true,"notes":"Lite tier only (NLI cross-encoder): 2,996 sentences took 658 s on 8 CPU threads (about 0.2 s each)."},{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code tool, not run here for this vertical."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services."}],"latency":[{"lane":"check of 4-5 judged sentences, hosted gateway route","typical_ms":1400,"source":"measured on our server 2026-09-24: 0.9-3.0 s for the four samples with a quiet gateway; 12-29 s earlier the same day while other evals saturated the shared gateway"},{"lane":"one judge call, 6 in flight, direct route","typical_ms":2000,"source":"measured on our server 2026-09-24: 1,995 calls in 679 s during the RAGTruth test run (GPU shared with other work)"},{"lane":"one sentence, NLI cross-encoder on CPU","typical_ms":220,"source":"measured on our server 2026-09-24: 2,996 sentences in 658 s, 8 threads"}],"benchmark":null,"notes":["Two operating points from one model. Counting 'partial' as flagged finds 97% of the sentences RAGTruth annotators marked but flags 30% of all sentences; many of those are added details the annotators left alone. So the default gate blocks only unsupported and contradicted (precision 0.59, recall 0.56 on the held-out test) and flags partial for a human look.","Confidence is per verdict, from the dev set: supported 99%, unsupported 60%, contradicted 42%, partial 20%. It is agreement with lenient human labels, not the probability that a sentence is false. The judge almost always says it is highly certain, so stated certainty adds little; a real per-sentence probability needs the verdict token's logprob, which the gateway does not return today.","The prompt was revised twice on the dev split; the test split was run once. The test run used the direct route to the same vLLM server because the shared gateway was saturated at the time.","Small source sets (up to 24,000 characters) go to the judge whole, sources first, so the model server reuses its prefix cache across the sentences of one check; larger sets are narrowed per sentence with BM25."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: text and sources live in memory for the request. The signed report carries hashes and character offsets, not your text."},{"label":"Inputs","value":"Text up to 8,000 characters and 40 sentences; up to 10 sources and 60,000 characters in total, as text (files are read in your browser on the console)."},{"label":"Typical hosted cost","value":"A fraction of a cent for a short answer against a one-page source at the gateway's list price (measured). Each run shows its own measured cost."},{"label":"What leaves the box (self-host)","value":"Nothing: the judge runs on your vLLM (or on CPU with the NLI option) and reports are signed with this box's key."}],"data_handling":{"page":"/data#grounding","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: text and sources live in memory for the request. The signed report carries hashes and character offsets, not your text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/developer/grounding","input":"grounding","lanes":[{"id":"sentences","title":"Sentence verdicts","kind":"list"},{"id":"gate","title":"Claims gate","kind":"markdown"},{"id":"report","title":"Signed report","kind":"json"}],"samples":[{"n":1,"id":"outreach-email","title":"Outreach email","deep_link":"/tools/developer/grounding?sample=1&autorun=0"},{"n":2,"id":"bakery-faq","title":"Bakery faq","deep_link":"/tools/developer/grounding?sample=2&autorun=0"},{"n":3,"id":"support-answer","title":"Support answer","deep_link":"/tools/developer/grounding?sample=3&autorun=0"},{"n":4,"id":"board-summary","title":"Board summary","deep_link":"/tools/developer/grounding?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/grounding-hosted.md","selfhost":"/prompts/grounding-selfhost.md","assemble":"/prompts/grounding-assemble.md","mac":"/prompts/grounding-mac.md"},"rehearsal":{"bundle":"/samples/grounding.zip","bundle_url":"https://decosa.ai/samples/grounding.zip","folder":"/samples/grounding/","expected":"/samples/grounding/expected.json","files":["/samples/grounding/expected.json","/samples/grounding/inputs/answer.txt","/samples/grounding/inputs/question.txt","/samples/grounding/inputs/sources.json"],"bytes":2105,"checks":["the claims gate blocks the answer","every claim sentence got a verdict (none errored)","the battery-life sentence is contradicted or not backed","the 240 V heater sentence is contradicted or not backed","the factory-reset sentence is supported","the signed report verifies against the same text and sources","the verified report matches the text","the verified report matches the sources","a report with one verdict changed no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: a fictional product (Halden T3) and manual written for Decosa, no real people or products. Part of decosa-api, AGPL-3.0-or-later.","about":"A support answer about a fictional thermostat, checked sentence by sentence against its manual. Two claims conflict with the manual (battery life, 240 V heaters), so the gate must block it, and the signed report must verify and catch a changed verdict.","run":{"containers":"docker compose exec api python scripts/rehearse.py grounding","checkout":"python scripts/rehearse.py grounding --bundle grounding.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py grounding"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=grounding","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1377.6,"basis":"estimate","unknown":[]},{"id":"alternate-fast-judge","gpu_gb":57,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/grounding","json":"/use-cases/grounding.json","metrics":"/metrics/grounding","console":"/tools/developer/grounding","console_sample":"/tools/developer/grounding?sample=1&autorun=0","stack":"/tools/developer/grounding#stack","try_live":"/tools/developer/grounding","watch":"/tools/developer/grounding","build":"/tools/developer/grounding#build","self_host":"/tools/developer/grounding#self-host","prompts":{"hosted":"/prompts/grounding-hosted.md","selfhost":"/prompts/grounding-selfhost.md","assemble":"/prompts/grounding-assemble.md","mac":"/prompts/grounding-mac.md"}}}