Skip to content
decosa

25 · Compliance and trust · live

Security questionnaire answerer

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Fill precision0.970 (64 of 66)test splitn = 66Dev: 0.963 (26 of 27). BM25 baseline on test: 0.606. With candidates from the evidence retrieval block (27 Sep, same test split): 0.984 (63 of 64).
  • Coverage (answerable questions filled correctly)0.984 (63 of 64)test splitn = 64BM25 baseline: 0.672.
  • Abstention (unanswerable questions sent to a person)0.969 (31 of 32)test splitn = 32Dev: 0.917 (11 of 12). BM25 baseline: 0.625. With the evidence retrieval block (27 Sep, same test split): 1.000 (32 of 32).
  • Stale approved answers caught4 of 4test splitn = 4
  • Filled answers with an invented claim0test splitn = 66Lexical check: 0 misses.

Dataset

A fictional vendor's library (50 approved answers, 9 documents; 2 answers deliberately out of date) and hand-written questions labelled fill, flag or stale before any model run: 40 dev, 100 test held out and run once after the prompts were frozen.

Caveats

  • Synthetic, single-author, one small library, English only: the same author wrote the library, the questions and the labels.
  • Question phrasing follows the library's vocabulary more closely than real buyer sheets; expect lower coverage and more review on real questionnaires.
  • Omissions are not checked: a trimmed answer can drop a qualifying sentence and still pass.
  • Only two stale entries, a very small sample.
  • "Supported" is a model judgement; the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
63 s
Receipts
62
Model calls
n/a
Tokens
n/a
Cost per run
$0.031

Self-host verification

Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers

Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. The 30-question sample: 21 approved, 2 review (TVM-01, LOG-01), 7 for a person, in 19 s with 61 calls; the XLSX export has a row per question with blank answers on the none rows; the record verifies and a changed status fails.

Rehearsal bundle: security-questionnaire.zip (7 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Speed depends on load: the 30-question sample took about 20 s on a quiet self-hosted GPU, about 60 s hosted, and about 2-3 minutes while the shared GPU was busy (25 Sep 2026).
  • PDFs and Word files are not read: send their text as documents.
  • The checks are model judgements and do not catch an answer that leaves out a qualifier; a person reads the sheet before it is sent. The eval library is synthetic.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checksAGPL-3.0-or-later
  • Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per questionQwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Apache-2.0 (both)
  • Selector (one call per question) and grounding judge (one call per changed or checked sentence)Qwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 32 GB card, no policy check (2)
  • Held-out test: fill precision without the source check: 0.928 (64 of 69)derived from docs/evals/security-questionnaire/test.json: the 3 answers the policy check sent to review would have been filled
  • Held-out test: coverage / abstention: 0.984 / 0.969 (unchanged: the policy check only affects stale answers)docs/evals/security-questionnaire.md
Standard · one GPU, selection plus both checks (hosted demo) (9)
  • Held-out test (100 questions): fill precision: 0.970 (64 of 66)docs/evals/security-questionnaire.md, synthetic library, labels written before any run
  • Held-out test: coverage of answerable questions: 0.984 (63 of 64)docs/evals/security-questionnaire.md
  • Held-out test: correct abstention when nothing approved fits: 0.969 (31 of 32)docs/evals/security-questionnaire.md
  • Held-out test: stale approved answers caught: 4 of 4docs/evals/security-questionnaire.md
  • Invented claims in filled answers (dev + test, 93 fills): 0; lexical check 0 missesdocs/evals/security-questionnaire.md
  • Baseline, BM25 top-1 with a threshold from dev: precision / coverage / abstention: 0.606 / 0.672 / 0.625docs/evals/security-questionnaire.md
  • With the evidence retrieval block (27 Sep, same held-out test, run once): precision / coverage / abstention / stale: 0.984 (63 of 64) / 0.984 (63 of 64) / 1.000 (32 of 32) / 4 of 4docs/evals/retrieval.md; runs in docs/evals/security-questionnaire/retrieval/
  • Prompt tokens for the 100 held-out questions, BM25 candidates vs the retrieval block: 218,203 vs 151,150 (-31%)docs/evals/retrieval.md
  • Retrieval alone (reranker top-1, threshold from dev, no model call): precision / coverage / abstention: 0.864 / 0.797 / 0.844 (BM25: 0.606 / 0.672 / 0.625)docs/evals/retrieval.md
Wanted · the largest open models, long context (1)
  • answer precision against the BM25 baseline, same protocol as standard: not measured yet

How we measure · All tools