{"schema_version":"1","site":"https://decosa.ai","id":"security-questionnaire","num":"25","name":"Security questionnaire answerer","tool_name":"Answer a security questionnaire","short":"Questionnaires","blurb":"Fills a buyer's security questionnaire from answers you have already approved. The model chooses and may trim, but anything it changes is checked against the source, stale answers are caught against your current policies, and questions with nothing approved behind them are left for a person.","status":"live","labels":{"industry":["compliance-trust"],"job":["draft","review"],"input":["files","text"],"deploy":["hosted","selfhost"],"status":"live","output":["data","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["compliance-trust"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["grounding","signed-record","evidence-retrieval"],"models":"Qwen3.8-27B","where":"Self-host (your policies stay on your machine) or hosted","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing, checks and export run on CPU","final_artifact":"The filled sheet (XLSX or CSV) with a source per answer, the questions for a person, and a signed review record.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per selection and check; signed review record","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":62952,"p95_ms":null,"runs":null,"receipts_per_run":62,"cost_per_run_usd":0.0313},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. The 30-question sample: 21 approved, 2 review (TVM-01, LOG-01), 7 for a person, in 19 s with 61 calls; the XLSX export has a row per question with blank answers on the none rows; the record verifies and a changed status fails."},"known_limits":["Speed depends on load: the 30-question sample took about 20 s on a quiet self-hosted GPU, about 60 s hosted, and about 2-3 minutes while the shared GPU was busy (25 Sep 2026).","PDFs and Word files are not read: send their text as documents.","The checks are model judgements and do not catch an answer that leaves out a qualifier; a person reads the sheet before it is sent. The eval library is synthetic."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Fill precision","value":"0.970 (64 of 66)","unit":null,"n":66,"split":"test","note":"Dev: 0.963 (26 of 27). BM25 baseline on test: 0.606. With candidates from the evidence retrieval block (27 Sep, same test split): 0.984 (63 of 64)."},{"name":"Coverage (answerable questions filled correctly)","value":"0.984 (63 of 64)","unit":null,"n":64,"split":"test","note":"BM25 baseline: 0.672."},{"name":"Abstention (unanswerable questions sent to a person)","value":"0.969 (31 of 32)","unit":null,"n":32,"split":"test","note":"Dev: 0.917 (11 of 12). BM25 baseline: 0.625. With the evidence retrieval block (27 Sep, same test split): 1.000 (32 of 32)."},{"name":"Stale approved answers caught","value":"4 of 4","unit":null,"n":4,"split":"test","note":null},{"name":"Filled answers with an invented claim","value":"0","unit":null,"n":66,"split":"test","note":"Lexical check: 0 misses."}],"dataset":"A fictional vendor's library (50 approved answers, 9 documents; 2 answers deliberately out of date) and hand-written questions labelled fill, flag or stale before any model run: 40 dev, 100 test held out and run once after the prompts were frozen.","held_out":true,"caveats":["Synthetic, single-author, one small library, English only: the same author wrote the library, the questions and the labels.","Question phrasing follows the library's vocabulary more closely than real buyer sheets; expect lower coverage and more review on real questionnaires.","Omissions are not checked: a trimmed answer can drop a qualifying sentence and still pass.","Only two stale entries, a very small sample.","\"Supported\" is a model judgement; the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/security-questionnaire"},"stack":{"summary":"Upload your approved answer library (past questionnaires, trust-center FAQ, policies, a SOC 2 summary) and the buyer's sheet (XLSX, CSV or plain text). For each question an open model chooses approved answers or policy passages, or says nothing fits. It may trim and join them but not add to them: any sentence it changes goes to the grounding judge, and if anything is added the answer reverts to the approved text word for word. Approved answers are also checked against the policies they cite, so stale ones are caught. You get the filled sheet as XLSX or CSV, a source and receipt for every answer, and a signed review record.","tagline":"Fills a buyer's security questionnaire from answers you have already approved, checks every edit against its source, and sends the rest to a person.","deployment":"hosted-or-self-host","regulatory_note":"Answers to a security questionnaire can become representations in a contract, so a person should read the sheet before it is sent; this tool fills only what your library already says and marks what needs sign-off, but the checks are model judgements and can be wrong. A trimmed answer can leave out a qualifying sentence, and the checks do not catch omissions. Your policies describe your security posture: self-host to keep them on your own machine. The hosted API keeps nothing (library and questionnaire stay in memory for the request; the review record holds hashes only; logs carry counts). Questionnaire templates carry their own licences: the CSA CAIQ is a free download from the Cloud Security Alliance after sign-in, but we found no licence that allows redistributing it, so the demo sheet is written from scratch in its shape. Not legal advice. Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.","components":[{"id":"answerer","role":"Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)","name":"decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12, standard library XLSX and CSV reader and writer, BM25 over approved answers and policy passages","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","wanted"],"alternative_to":null},{"id":"retrieval","role":"Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per question","name":"Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","hf_repo":"Qwen/Qwen3-Reranker-4B","license":"Apache-2.0 (both)","params":"0.6B + 4B","quant":"BF16","vram_gb":12,"memory_gb_estimate":null,"engine":"transformers 5.17 in services/retrieval, one process on the same GPU; revisions 97b0c61 and 22e6836","receipt_coverage":"partial","in_hosted_demo":false,"tiers":["standard"],"alternative_to":null},{"id":"llm","role":"Selector (one call per question) and grounding judge (one call per changed or checked sentence)","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off, prefix caching","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"gemma-selector","role":"Fast selector with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0","params":"26B","quant":"BF16 (49 GB)","vram_gb":49,"memory_gb_estimate":null,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null},{"id":"dsv41-wanted","role":"Selector and drafter with long context","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"glm-wanted","role":"Checker, from another family","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 32 GB card, no policy check","summary":"The same selector and edit check with the source check off (source_check: false): about half the model calls, but stale approved answers are filled instead of flagged.","components":["answerer","llm"],"hardware":"1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Held-out test: fill precision without the source check","value":"0.928 (64 of 69)","source":"derived from docs/evals/security-questionnaire/test.json: the 3 answers the policy check sent to review would have been filled"},{"metric":"Held-out test: coverage / abstention","value":"0.984 / 0.969 (unchanged: the policy check only affects stale answers)","source":"docs/evals/security-questionnaire.md"}],"latency_note":"estimate: fewer calls per question than standard; not measured on a 32 GB card.","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Self-hosted: every call is signed by the box's own key (an attestation), with no gateway countersign.","hosting":null},{"id":"standard","label":"Standard · one GPU, selection plus both checks (hosted demo)","summary":"Qwen3.8-27B chooses the answers and judges every changed sentence and every cited policy, each in its own receipted call. This is what the hosted API runs.","components":["answerer","retrieval","llm"],"hardware":"1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Held-out test (100 questions): fill precision","value":"0.970 (64 of 66)","source":"docs/evals/security-questionnaire.md, synthetic library, labels written before any run"},{"metric":"Held-out test: coverage of answerable questions","value":"0.984 (63 of 64)","source":"docs/evals/security-questionnaire.md"},{"metric":"Held-out test: correct abstention when nothing approved fits","value":"0.969 (31 of 32)","source":"docs/evals/security-questionnaire.md"},{"metric":"Held-out test: stale approved answers caught","value":"4 of 4","source":"docs/evals/security-questionnaire.md"},{"metric":"Invented claims in filled answers (dev + test, 93 fills)","value":"0; lexical check 0 misses","source":"docs/evals/security-questionnaire.md"},{"metric":"Baseline, BM25 top-1 with a threshold from dev: precision / coverage / abstention","value":"0.606 / 0.672 / 0.625","source":"docs/evals/security-questionnaire.md"},{"metric":"With the evidence retrieval block (27 Sep, same held-out test, run once): precision / coverage / abstention / stale","value":"0.984 (63 of 64) / 0.984 (63 of 64) / 1.000 (32 of 32) / 4 of 4","source":"docs/evals/retrieval.md; runs in docs/evals/security-questionnaire/retrieval/"},{"metric":"Prompt tokens for the 100 held-out questions, BM25 candidates vs the retrieval block","value":"218,203 vs 151,150 (-31%)","source":"docs/evals/retrieval.md"},{"metric":"Retrieval alone (reranker top-1, threshold from dev, no model call): precision / coverage / abstention","value":"0.864 / 0.797 / 0.844 (BM25: 0.606 / 0.672 / 0.625)","source":"docs/evals/retrieval.md"}],"latency_note":"measured: a couple of minutes for the sample sheet through the shared gateway; answers stream in as they finish.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every selection and judgement is a separate gateway call with a gateway-signed receipt; the signed record lists them all.","hosting":null},{"id":"wanted","label":"Wanted · the largest open models, long context","summary":"DeepSeek-V4.1-Flash selects and drafts answers with much more of the evidence library in context, and GLM-5.3-Flash checks each one. Both models support 1M tokens (vendor); the proposed network need serves 131k. Not served yet.","components":["answerer","llm","dsv41-wanted","glm-wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.","quality_evidence":[{"metric":"answer precision against the BM25 baseline, same protocol as standard","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[{"id":"fast-judge","label":"Fast option-token selector","components":["gemma-selector"],"hardware":"1x RTX PRO 6000 96 GB (BF16 weights are 49 GB)","use":"A 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /questionnaire/info, /questionnaire/samples; POST /questionnaire/parse, /questionnaire/answer (SSE or JSON), /questionnaire/export (XLSX or CSV), /questionnaire/verify. Keeps no text."},{"name":"vLLM (selector and judge)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"Synthetic eval set (docs/evals/security-questionnaire)","url":null,"license":"Apache-2.0","purpose":"A fictional vendor's library (50 approved answers, 9 policy and report documents, 2 answers deliberately out of date) and 140 hand-labelled buyer questions: 40 dev, 100 held-out test, written before any model run."},{"name":"Grounding check (tool 17)","url":null,"license":"Apache-2.0","purpose":"The per-sentence judge used for both checks: edited answers against the chosen text, and approved answers against the policies they cite."},{"name":"POST /questionnaire/verify","url":null,"license":"Apache-2.0","purpose":"Checks a review record's signature against this server's key and, if you send them, the answer hashes. The console also checks the signature in your browser with WebCrypto."}],"hardware":[{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; each call is about 2,200 prompt tokens. Estimate: not run here for this tool."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services."}],"latency":[{"lane":"30-question sheet, hosted gateway route, 6 questions in flight","typical_ms":108600,"source":"measured on our server 2026-09-25: the Corvid Freight demo sheet, 61 receipted calls, answers stream in as each question finishes"},{"lane":"100-question eval sheet, gateway route, 4 in flight","typical_ms":548000,"source":"measured on our server 2026-09-25: 177 calls, about 2,200 prompt and 120 generated tokens per question (GPU shared with other work)"}],"benchmark":null,"notes":["Select, don't generate. Of 66 answers filled on the held-out test, 22 were the approved text word for word, 29 were whole sentences of it (checked in code, no model), and 15 were edits the grounding judge found supported. The judge rejected no edit in 93 real fills; the revert path is covered by unit tests.","Nothing approved, nothing drafted: 31 of 32 test questions with no approved answer went to a person. A top-1 retrieval baseline fills a third of them with the nearest answer.","Stale answers: approved answers are checked against the policies they cite. Both out-of-date answers in the demo library were flagged every time they were chosen (4 of 4 on test).","The eval is synthetic, written by one author, on one small library, so read it as a check of the mechanism rather than accuracy on real questionnaires.","Formats: XLSX (first sheet with a Question header; ID and domain columns detected), CSV and plain text in; XLSX (no formulas) and CSV (formula-like cells neutralised) out. PDFs and Word files are not read: paste their text."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: the library and the questionnaire live in memory for the request. The review record holds hashes, source ids and statuses, never text."},{"label":"What leaves the box","value":"Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box. A library describes your security posture: for real policies, self-host."},{"label":"Input formats","value":"Questionnaire as XLSX (first sheet with a Question header), CSV, pasted text or JSON; library as approved answers (CSV, XLSX or JSON) plus policy documents as text. Out: XLSX or CSV. Up to 60 questions per demo run, 200 with a key."},{"label":"Typical run","value":"The sample sheet against its library: a few dozen model calls, a few cents at the gateway list price. Each run shows its own measured cost."}],"data_handling":{"page":"/data#security-questionnaire","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: the library and the questionnaire live in memory for the request. The review record holds hashes, source ids and statuses, never text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/finance/security-questionnaire","input":"questionnaire","lanes":[{"id":"answers","title":"Answers","kind":"list"},{"id":"review","title":"For a person","kind":"list"},{"id":"record","title":"Signed review record","kind":"json"}],"samples":[{"n":1,"id":"corvid-freight","title":"Corvid freight","deep_link":"/tools/finance/security-questionnaire?sample=1&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/security-questionnaire-hosted.md","selfhost":"/prompts/security-questionnaire-selfhost.md","assemble":"/prompts/security-questionnaire-assemble.md","mac":"/prompts/security-questionnaire-mac.md"},"rehearsal":{"bundle":"/samples/security-questionnaire.zip","bundle_url":"https://decosa.ai/samples/security-questionnaire.zip","folder":"/samples/security-questionnaire/","expected":"/samples/security-questionnaire/expected.json","files":["/samples/security-questionnaire/expected.json","/samples/security-questionnaire/inputs/library.json","/samples/security-questionnaire/inputs/questionnaire.csv"],"bytes":6795,"checks":["the CSV questionnaire parses into four questions","the security-policy question (GOV-01) fills from an approved answer","the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2","the out-of-date pen-test answer (TVM-01) is not passed as approved","the bug-bounty question (TVM-03) is left for a person","no question errored","the signed review record verifies","the answers match the hashes in the record","a record with the bug-bounty row forged to approved no longer verifies","every model call has a signed receipt"],"licence":"Fictional: Tallyloom, Corvid Freight and Pellham & Vose LLP are invented. The questionnaire is written for Decosa in the shape of the CSA CAIQ (domain, question id, question); it is not CAIQ text. Part of decosa-api, AGPL-3.0-or-later.","about":"Four rows of a fictional buyer's security questionnaire (CSV), answered from a fictional vendor's approved answer library and policies. The policy and encryption questions must fill from approved answers, the out-of-date pen-test answer must not pass as approved, the bug-bounty question (nothing approved) must be left for a person, and the signed review record must verify and catch a forged status.","run":{"containers":"docker compose exec api python scripts/rehearse.py security-questionnaire","checkout":"python scripts/rehearse.py security-questionnaire --bundle security-questionnaire.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py security-questionnaire"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=security-questionnaire","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":69.6,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1377.6,"basis":"estimate","unknown":[]},{"id":"alternate-fast-judge","gpu_gb":57,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/finance/security-questionnaire","json":"/use-cases/security-questionnaire.json","metrics":"/metrics/security-questionnaire","console":"/tools/finance/security-questionnaire","console_sample":"/tools/finance/security-questionnaire?sample=1&autorun=0","stack":"/tools/finance/security-questionnaire#stack","try_live":"/tools/finance/security-questionnaire","watch":"/tools/finance/security-questionnaire","build":"/tools/finance/security-questionnaire#build","self_host":"/tools/finance/security-questionnaire#self-host","prompts":{"hosted":"/prompts/security-questionnaire-hosted.md","selfhost":"/prompts/security-questionnaire-selfhost.md","assemble":"/prompts/security-questionnaire-assemble.md","mac":"/prompts/security-questionnaire-mac.md"}}}