Security questionnaire answerer: eval on a synthetic library (25 Sep 2026)
Questions: does it pick the right approved answer, does it hold back when nothing approved answers the question, does it catch approved answers that have gone stale, and does anything it fills say more than its source?
Data
- Library: the fictional vendor Tallyloom (
decosa_api/verticals/questionnaire/data/samples.json, built byscripts/questionnaire_library.py): 50 approved answers (49 cite the policy they rest on) and 9 documents (policies, an incident response summary and a SOC 2 summary, a few paragraphs each). Two approved answers are deliberately out of date against their policy (pen tests "once a year" vs "at least twice a year"; logs kept "90 days" vs "1 year"). - Questions:
docs/evals/security-questionnaire/eval.json, written by hand on 25 Sep 2026 before any model run, in varied buyer phrasing (yes/no, "describe", "state", compound questions). Each is labelled:fill: an approved answer or a policy passage answers it fully; the gold is the set of acceptable sources (entry ids, or a document id for a passage from that document);flag: nothing approved answers it (FedRAMP, bug bounty, DLP, PAM, SOC 1, uptime SLA...), including near misses on a covered topic (a 90-day key rotation, a per-customer HSM, the cyber-insurance amount);stale: it hits one of the two out-of-date answers.
- Dev: 40 questions (26 fill, 12 flag, 2 stale). Test: 100 questions (64 fill, 32 flag, 4 stale), held out: run once, after the prompts were frozen. The dev run needed no prompt change, so the prompts are the first version.
- The same author wrote the library, the questions and the labels, and the phrasing follows the library's vocabulary more closely than a real buyer's sheet would. Treat these numbers as a check that the mechanism works, not as the accuracy on real questionnaires.
Metrics
- Fill precision: of the questions the system filled (status
approvedorpolicy), the share whose every chosen source is in the gold set. A fill on aflagquestion counts as wrong; a fill of a stale answer counts as wrong. - Coverage:
fillquestions filled correctly / allfillquestions. - Abstention:
flagquestions sent to a person (statusnoneorreview) / allflagquestions. - Stale caught: stale questions flagged, or filled only from the current policy.
- Invented claims: filled answers with a claim sentence that is neither a whole sentence of its sources nor judged
supported, plus an independent lexical check: every number in the answer must be in its sources and at least 85% of its content words.
Results
Hosted route (model gateway, Qwen3.8-27B NVFP4, temperature 0, thinking off; every call gateway-receipted).
| Split | Fill precision | Coverage | Abstention | Stale caught | Invented claims | Filled answers: verbatim / trimmed / edited and grounded / reverted |
|---|---|---|---|---|---|---|
| Test (100, held out) | 0.970 (64 of 66) | 0.984 (63 of 64) | 0.969 (31 of 32) | 4 of 4 | 0 (lexical check: 0 misses) | 22 / 29 / 15 / 0 |
| Dev (40) | 0.963 (26 of 27) | 1.000 (26 of 26) | 0.917 (11 of 12) | 2 of 2 | 0 (0 misses) | 11 / 12 / 4 / 0 |
| Baseline, test: BM25 top-1 approved answer, filled above a score threshold chosen on dev (6.5) | 0.606 | 0.672 | 0.625 | 0 of 4 | n/a (copies) |
Errors on test:
- t084 "Are keys stored in a hardware security module dedicated to each customer?" was filled with the BYOK and AWS KMS
answers. Nothing in it is false, but it does not answer the question; labelled
flag. - t056 "Do laptops lock automatically when idle, and after how long?" was answered from the endpoint answer (AL-020) plus the policy passage with the 10-minute figure. The gold listed only the policy; AL-020 is arguably fine, but it is scored as an error.
- Dev d35 "Do you rotate encryption keys every 90 days?" was answered with "rotated automatically every year". A
reviewer might accept that; the label says
flag, and we left the label alone.
Also measured on the hosted demo sample (30 questions): 19 approved, 2 policy, 2 review (both stale answers caught), 7 sent to a person, 61 receipted model calls, 108.6 s end to end with 6 questions in flight. One of the seven, "How do you assess the security of your subprocessors?", has an approved answer about vendors (AL-033) that a reviewer would use: the selector said NONE. Holding back like that is the intended failure direction, but it costs time.
Reading the numbers
- Nothing invented. No filled answer contains a sentence outside its sources. That is partly by construction: 51 of 66 test answers are the approved text or whole sentences of it (checked in code, no model), and the other 15 were edits that the grounding judge found supported. Reverts (the judge rejecting an edit) happened 0 times in 93 real fills: the selector nearly always copies. The revert path is covered by unit tests, not by this eval.
- Abstention is the point. A plain retrieval baseline fills a third of the unanswerable questions with the nearest approved answer. The selector sends 31 of 32 to a person.
- Stale answers. The source check found both out-of-date answers every time they were chosen (4 of 4 on test, 2 of 2 on dev, 2 of 2 in the demo run). Two stale entries is a very small sample.
- Omissions are not checked. A trimmed answer can drop a qualifying sentence ("on the Enterprise plan only") and every remaining sentence still passes. The selector is told to drop only sentences that do not bear on the question; a person should still read trimmed answers. The Checks column in the export says which answers were trimmed.
Honest caveats
- Synthetic, single-author, one small library, English only. Real libraries have hundreds of overlapping answers of
mixed age and quality, and real questions are longer and vaguer; expect lower coverage and more
review. - "Supported" is a model judgement (the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth at its
default gate; see
docs/evals/grounding.md). Here it is used on short, near-copied answers, which is its easiest case. - The labels were not changed after seeing results. Two scored errors (t056, d35) are arguable, as noted.
Reproduce
python scripts/questionnaire_library.py # rebuilds samples.json
python scripts/questionnaire_eval.py run --split dev --work ~/.cache/questionnaire-eval/v1
python scripts/questionnaire_eval.py run --split test --work ~/.cache/questionnaire-eval/v1
python scripts/questionnaire_eval.py score --split test --work ~/.cache/questionnaire-eval/v1
Runs and scores are in docs/evals/security-questionnaire/{dev,test}.json and {dev,test}-score.json.
Cost on our server (gateway route, GPU1 shared): dev 64 calls in 161 s; test 177 calls in 548 s, 218 k prompt tokens and
11.7 k generated (about 2,200 prompt and 120 generated tokens per question), 4 questions in flight.
27 Sep 2026: candidates from the evidence retrieval block
- What changed: with
DECOSA_RETRIEVAL_URLset, the candidates come from the retrieval block: Qwen3-Embedding-0.6B + BM25, reranked by Qwen3-Reranker-4B, 4 approved answers + 2 policy passages (BM25 gave 6 + 3). Nothing else changed. - The run: the test split was run once more, after a dev run that matched the old dev numbers exactly.
| Held-out test | BM25 candidates (above) | Retrieval block |
|---|---|---|
| Fill precision | 0.970 (64 of 66) | 0.984 (63 of 64) |
| Coverage | 0.984 (63 of 64) | 0.984 (63 of 64) |
| Abstention | 0.969 (31 of 32) | 1.000 (32 of 32) |
| Stale caught | 4 of 4 | 4 of 4 |
| Invented claims | 0 | 0 |
| Prompt tokens (177 vs 172 calls) | 218,203 | 151,150 (−31%) |
- What moved: t084 (the per-customer HSM question) is now sent to a person. t056 is the one remaining error, as before.
- Receipts: each question's search has a signed receipt, and the review record adds
retrieval(index hash, model revisions, search receipt ids). - Retrieval alone, no model call: the reranker's top answer, filled above a threshold chosen on dev, scores 0.864 precision against BM25's 0.606. That closes 71% of the gap to the selector, but it catches only 1 of 4 stale answers.
- Where it is written up: details and caveats are in
docs/evals/retrieval.md; the runs are indocs/evals/security-questionnaire/retrieval/.