25 · Compliance and trust · live
Security questionnaire answerer
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Fill precision0.970 (64 of 66)test splitn = 66Dev: 0.963 (26 of 27). BM25 baseline on test: 0.606. With candidates from the evidence retrieval block (27 Sep, same test split): 0.984 (63 of 64).
- Coverage (answerable questions filled correctly)0.984 (63 of 64)test splitn = 64BM25 baseline: 0.672.
- Abstention (unanswerable questions sent to a person)0.969 (31 of 32)test splitn = 32Dev: 0.917 (11 of 12). BM25 baseline: 0.625. With the evidence retrieval block (27 Sep, same test split): 1.000 (32 of 32).
- Stale approved answers caught4 of 4test splitn = 4
- Filled answers with an invented claim0test splitn = 66Lexical check: 0 misses.
Dataset
A fictional vendor's library (50 approved answers, 9 documents; 2 answers deliberately out of date) and hand-written questions labelled fill, flag or stale before any model run: 40 dev, 100 test held out and run once after the prompts were frozen.
Caveats
- Synthetic, single-author, one small library, English only: the same author wrote the library, the questions and the labels.
- Question phrasing follows the library's vocabulary more closely than real buyer sheets; expect lower coverage and more review on real questionnaires.
- Omissions are not checked: a trimmed answer can drop a qualifying sentence and still pass.
- Only two stale entries, a very small sample.
- "Supported" is a model judgement; the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 63 s
- Receipts
- 62
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.031
Self-host verification
Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. The 30-question sample: 21 approved, 2 review (TVM-01, LOG-01), 7 for a person, in 19 s with 61 calls; the XLSX export has a row per question with blank answers on the none rows; the record verifies and a changed status fails.
Rehearsal bundle: security-questionnaire.zip (7 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Speed depends on load: the 30-question sample took about 20 s on a quiet self-hosted GPU, about 60 s hosted, and about 2-3 minutes while the shared GPU was busy (25 Sep 2026).
- PDFs and Word files are not read: send their text as documents.
- The checks are model judgements and do not catch an answer that leaves out a qualifier; a person reads the sheet before it is sent. The eval library is synthetic.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checksAGPL-3.0-or-later
- Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per questionQwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Apache-2.0 (both)
- Selector (one call per question) and grounding judge (one call per changed or checked sentence)Qwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 32 GB card, no policy check (2)
- Held-out test: fill precision without the source check: 0.928 (64 of 69)derived from docs/evals/security-questionnaire/test.json: the 3 answers the policy check sent to review would have been filled
- Held-out test: coverage / abstention: 0.984 / 0.969 (unchanged: the policy check only affects stale answers)docs/evals/security-questionnaire.md
Standard · one GPU, selection plus both checks (hosted demo) (9)
- Held-out test (100 questions): fill precision: 0.970 (64 of 66)docs/evals/security-questionnaire.md, synthetic library, labels written before any run
- Held-out test: coverage of answerable questions: 0.984 (63 of 64)docs/evals/security-questionnaire.md
- Held-out test: correct abstention when nothing approved fits: 0.969 (31 of 32)docs/evals/security-questionnaire.md
- Held-out test: stale approved answers caught: 4 of 4docs/evals/security-questionnaire.md
- Invented claims in filled answers (dev + test, 93 fills): 0; lexical check 0 missesdocs/evals/security-questionnaire.md
- Baseline, BM25 top-1 with a threshold from dev: precision / coverage / abstention: 0.606 / 0.672 / 0.625docs/evals/security-questionnaire.md
- With the evidence retrieval block (27 Sep, same held-out test, run once): precision / coverage / abstention / stale: 0.984 (63 of 64) / 0.984 (63 of 64) / 1.000 (32 of 32) / 4 of 4docs/evals/retrieval.md; runs in docs/evals/security-questionnaire/retrieval/
- Prompt tokens for the 100 held-out questions, BM25 candidates vs the retrieval block: 218,203 vs 151,150 (-31%)docs/evals/retrieval.md
- Retrieval alone (reranker top-1, threshold from dev, no model call): precision / coverage / abstention: 0.864 / 0.797 / 0.844 (BM25: 0.606 / 0.672 / 0.625)docs/evals/retrieval.md
Wanted · the largest open models, long context (1)
- answer precision against the BM25 baseline, same protocol as standard: not measured yet