Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Answer a security questionnaire

A filled questionnaire built only from your approved answers, with the source of each one, gaps sent to a person, and stale answers flagged.

Held-out test0.970 (64 of 66)Filled answers that were correct (held-out test)
On production63 smedian on production (2026-09-25); slower when the service is busy
List price~$0.12 per 100 questionsmeasured, at list price

Built on: Grounding, Signed record, Evidence retrieval

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing, checks and export run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the security questionnaire answerer API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
security-questionnaire

Use the hosted API

# Decosa security questionnaire answerer: use the hosted API

You are wiring Decosa's questionnaire answerer into this project. It takes a company's approved answer library and a
buyer's security questionnaire and fills each question with approved text only: an open model chooses approved answers
or policy passages, any sentence it changes is checked against the source (and reverted if anything was added), approved
answers are checked against the policies they cite, and questions with nothing approved behind them are left for a
person. Every model call has its own signed receipt, and each run ends with a review record signed by the server. Use
only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API keeps nothing, but a library describes a company's security posture: for real policies, self-host.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`. Up to 200 questions a run.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "security-questionnaire"}` returns
   `{"token", "expires_at", "budget"}`. a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`); up to 60
   questions a run. A run needs 250 tokens of budget per question to start (402 otherwise). One run at a time per token.

## Inputs
- `questionnaire`: `{"questions": [{"id", "section", "question"}]}`, or `{"file": {"name": "buyer.xlsx", "b64": "..."}}`
  (XLSX or CSV: the first sheet with a header containing "question"), or `{"text": "one question per line"}`.
- `library`: `{"entries": [{"id", "question", "answer", "yes_no"?, "sources"?: ["POL-ENC"], "origin"?}],
  "documents": [{"id": "POL-ENC", "title", "kind": "policy"|"report"|"faq", "text"}], "files"?: [{"name", "b64"}]}`.
  A CSV or XLSX file in `files` needs Question and Answer columns (optional ID, Yes/No, Source, Origin).
- Limits: 1,500 entries, 25 documents (400,000 characters in total), 3 MB per file, 8 MB per request.
- `POST /questionnaire/parse` (token) with either input returns what will be used, with no model call; it lists
  citations that point at documents you did not send.

## Run
`POST /questionnaire/answer` (token) with `{"questionnaire": ..., "library": ...}`.
- JSON by default: `{counts, items: [{i, id, section, question, status, answer, yes_no, sources: [{ref, kind, label, text}],
  fit, reason, adaptation: {mode, proposed?, why?}, source_check: {state, docs} | null, receipt_ids}], receipts, record, budget, note}`.
- With `Accept: text/event-stream` it streams `ready`, then `receipt` and `item` events as questions finish (not in
  sheet order), then `record`, `budget`, `done`.
- Statuses: `approved` (approved answers), `policy` (built from a policy passage; needs sign-off), `review` (partial fit,
  or the approved answer conflicts with the policy it cites), `none` (nothing approved; `answer` is null), `error`.
- `adaptation.mode`: `verbatim`, `extract` (whole sentences of the approved text), `grounded` (an edit the judge found
  supported) or `reverted` (the edit added something; the approved text was used). `yes_no` is filled only when the
  approved answer gives the same Yes or No.

## Example: fill a sheet and save it (Python, `pip install httpx`)
```python
import base64, httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
b64 = lambda p: base64.b64encode(open(p, "rb").read()).decode()
body = {"questionnaire": {"file": {"name": "buyer.xlsx", "b64": b64("buyer.xlsx")}},
        "library": {"files": [{"name": "approved-answers.csv", "b64": b64("approved-answers.csv")}],
                    "documents": [{"id": "POL-ISP", "title": "Information Security Policy", "text": open("pol-isp.txt").read()}]}}
r = httpx.post(f"{API}/questionnaire/answer", json=body, headers=H, timeout=900)
r.raise_for_status()
run = r.json()
for it in run["items"]:
    if it["status"] in ("none", "review", "error"):
        print(it["id"], it["status"], it["reason"])
x = httpx.post(f"{API}/questionnaire/export", json={"items": run["items"], "format": "xlsx", "record": run["record"]}, headers=H, timeout=60)
open("buyer-filled.xlsx", "wb").write(x.content)
```

## Verify the review record (`pip install cryptography`)
```python
import json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
rec = run["record"]
pub = json.load(urllib.request.urlopen("https://api.decosa.ai/attest/signing-key"))["pubkey"]
assert rec["signer"] == pub
body = {k: v for k, v in rec.items() if k != "sig"}
msg = rec["v"].encode() + b"\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(rec["sig"]), msg)
```
Or `POST https://api.decosa.ai/questionnaire/verify` with `{"record": rec, "answers": [{"i", "answer"}]}`. The record holds
hashes, source ids and statuses, never text. Each id in `rec["receipts"]` resolves at `GET https://api.decosa.ai/receipts/{id}`.

## Honest limits
- A person reads the sheet before it is sent. The checks are model judgements, and they do not catch an answer that
  leaves out a qualifier. The eval numbers on the Stack tab come from a synthetic library.
- PDFs and Word files are not read: send their text as documents.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa security questionnaire answerer: run it yourself (containers)

You are setting up the Decosa questionnaire answerer on this machine, so our policies and approved answers never leave
it. It fills a buyer's security questionnaire from approved text only, checks every change against its source, and
signs a review record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/security-questionnaire.zip (7 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py security-questionnaire` (the api image carries the same bundle under /app/rehearsal/security-questionnaire/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py security-questionnaire --bundle security-questionnaire.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the CSV questionnaire parses into four questions", "the security-policy question (GOV-01) fills from an approved answer", "the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with 32 GB or more.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
   127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/questionnaire/info` lists the statuses and prompt hashes;
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify our records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"security-questionnaire"}`, fetch
   `GET /questionnaire/samples`, and send `samples[0].questionnaire` and `samples[0].library` to
   `POST /questionnaire/answer`. Expect most questions `approved`, TVM-01 and LOG-01 `review` (their approved answers
   conflict with the current policy), and FedRAMP, bug bounty and similar questions `none`. Then
   `POST /questionnaire/verify` with the record: `valid_signature` and `signed_by_this_server` must both be true.
6. Report back: the public key, the status counts, and how long the run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds our
policies. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/security-questionnaire-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) needs a GPU.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 32 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (69.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (69.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"security-questionnaire"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py security-questionnaire

Download the mock-data bundle (7 KB, 10 checks)expected.json

Four rows of a fictional buyer's security questionnaire (CSV), answered from a fictional vendor's approved answer library and policies. The policy and encryption questions must fill from approved answers, the out-of-date pen-test answer must not pass as approved, the bug-bounty question (nothing approved) must be left for a person, and the signed review record must verify and catch a forged status.

What the rehearsal checks
  • the CSV questionnaire parses into four questions
  • the security-policy question (GOV-01) fills from an approved answer
  • the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2
  • the out-of-date pen-test answer (TVM-01) is not passed as approved
  • the bug-bounty question (TVM-03) is left for a person
  • no question errored
  • the signed review record verifies
  • the answers match the hashes in the record
  • a record with the bug-bounty row forged to approved no longer verifies
  • every model call has a signed receipt

Licence: Fictional: Tallyloom, Corvid Freight and Pellham & Vose LLP are invented. The questionnaire is written for Decosa in the shape of the CSA CAIQ (domain, question id, question); it is not CAIQ text. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa security questionnaire answerer: run it yourself (containers)

You are setting up the Decosa questionnaire answerer on this machine, so our policies and approved answers never leave
it. It fills a buyer's security questionnaire from approved text only, checks every change against its source, and
signs a review record. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/security-questionnaire.zip (7 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py security-questionnaire` (the api image carries the same bundle under /app/rehearsal/security-questionnaire/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py security-questionnaire --bundle security-questionnaire.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the CSV questionnaire parses into four questions", "the security-policy question (GOV-01) fills from an approved answer", "the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with 32 GB or more.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
   127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/questionnaire/info` lists the statuses and prompt hashes;
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify our records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"security-questionnaire"}`, fetch
   `GET /questionnaire/samples`, and send `samples[0].questionnaire` and `samples[0].library` to
   `POST /questionnaire/answer`. Expect most questions `approved`, TVM-01 and LOG-01 `review` (their approved answers
   conflict with the current policy), and FedRAMP, bug bounty and similar questions `none`. Then
   `POST /questionnaire/verify` with the record: `valid_signature` and `signed_by_this_server` must both be true.
6. Report back: the public key, the status counts, and how long the run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds our
policies. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/security-questionnaire-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierSecurity questionnaire answerer on GeForce RTX 5090: use the Lite · one 32 GB card, no policy check tier

The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; each call is about 2,200 prompt tokens. Estimate: not run here for this use case.

Lite · one 32 GB card, no policy check: what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Answerer: decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks. CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Selector: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Security questionnaire answerer, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Security questionnaire answerer on my hardware

Fetch https://decosa.ai/prompts/security-questionnaire-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=security-questionnaire)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · one 32 GB card, no policy check (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Answerer: decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks, CPU
- Selector: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/security-questionnaire-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Security questionnaire answerer: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Security questionnaire answerer on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/security-questionnaire.zip (7 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py security-questionnaire` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the CSV questionnaire parses into four questions", "the security-policy question (GOV-01) fills from an approved answer", "the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per question | transformers on CUDA (services/retrieval), about 12 GB | The same service with PyTorch on MPS (bf16): same scores within rounding, about 12 GB, rerank of 40 passages 5.1 s (p50) vs 1.8 s on the GPU | Runs, measured |
| Selector (one call per question) and grounding judge (one call per changed or checked sentence) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py security-questionnaire`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 63 s · ~$0.031 per run · 62 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone, compose up, sample against local model servers

Measured cost to run: about $0.12 per 100 questions (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. The 30-question sample: 21 approved, 2 review (TVM-01, LOG-01), 7 for a person, in 19 s with 61 calls; the XLSX export has a row per question with blank answers on the none rows; the record verifies and a changed status fails.

Known limits (3)
  • Speed depends on load: the 30-question sample took about 20 s on a quiet self-hosted GPU, about 60 s hosted, and about 2-3 minutes while the shared GPU was busy (25 Sep 2026).
  • PDFs and Word files are not read: send their text as documents.
  • The checks are model judgements and do not catch an answer that leaves out a qualifier; a person reads the sheet before it is sent. The eval library is synthetic.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing, checks and export run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the security questionnaire answerer API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

Fills a buyer's security questionnaire from answers you have already approved, checks every edit against its source, and sends the rest to a person.

Upload your approved answer library (past questionnaires, trust-center FAQ, policies, a SOC 2 summary) and the buyer's sheet (XLSX, CSV or plain text). For each question an open model chooses approved answers or policy passages, or says nothing fits. It may trim and join them but not add to them: any sentence it changes goes to the grounding judge, and if anything is added the answer reverts to the approved text word for word. Approved answers are also checked against the policies they cite, so stale ones are caught. You get the filled sheet as XLSX or CSV, a source and receipt for every answer, and a signed review record.

Deployment
Hosted or self-host
Regulatory
Answers to a security questionnaire can become representations in a contract, so a person should read the sheet before it is sent; this tool fills only what your library already says and marks what needs sign-off, but the checks are model judgements and can be wrong. A trimmed answer can leave out a qualifying sentence, and the checks do not catch omissions. Your policies describe your security posture: self-host to keep them on your own machine. The hosted API keeps nothing (library and questionnaire stay in memory for the request; the review record holds hashes only; logs carry counts). Questionnaire templates carry their own licences: the CSA CAIQ is a free download from the Cloud Security Alliance after sign-in, but we found no licence that allows redistributing it, so the demo sheet is written from scratch in its shape. Not legal advice. Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.
Architecture
Text description

The buyer's questionnaire (XLSX, CSV or text) and the approved library (answers, policies, a SOC 2 summary) go to the answerer, which parses them and retrieves candidate answers and policy passages for each question with BM25. The Qwen3.8-27B selector chooses candidates or NONE, and says whether they answer the whole question. Sentences copied word for word pass a code check; changed sentences go to the grounding judge, and an answer with anything added reverts to the approved text. Approved answers that cite a policy are checked against it, and conflicts go to review. Outputs: the filled sheet with a status, source and receipts per answer, questions for a person, and a signed review record. On the hosted route each model call gets a signed receipt that our gateway countersigns. In self-host mode everything runs on your machine.

Architecture

At a glance

Data retention
Nothing stored: the library and the questionnaire live in memory for the request. The review record holds hashes, source ids and statuses, never text.
What leaves the box
Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box. A library describes your security posture: for real policies, self-host.
Input formats
Questionnaire as XLSX (first sheet with a Question header), CSV, pasted text or JSON; library as approved answers (CSV, XLSX or JSON) plus policy documents as text. Out: XLSX or CSV. Up to 60 questions per demo run, 200 with a key.
Typical run
The sample sheet against its library: a few dozen model calls, a few cents at the gateway list price. Each run shows its own measured cost.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 32 GB card, no policy check

    The same selector and edit check with the source check off (source_check: false): about half the model calls, but stale approved answers are filled instead of flagged.

    Models
    • decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Held-out test: fill precision without the source check0.928 (64 of 69)derived from docs/evals/security-questionnaire/test.json: the 3 answers the policy check sent to review would have been filled
    • Held-out test: coverage / abstention0.984 / 0.969 (unchanged: the policy check only affects stale answers)docs/evals/security-questionnaire.md
    Latency
    estimate: fewer calls per question than standard; not measured on a 32 GB card.
    Verification
    Proof: partialSelf-host onlySelf-hosted: every call is signed by the box's own key (an attestation), with no gateway countersign.
  • In the hosted demo

    Standard

    one GPU, selection plus both checks (hosted demo)

    Qwen3.8-27B chooses the answers and judges every changed sentence and every cited policy, each in its own receipted call. This is what the hosted API runs.

    Models
    • decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks
    • Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Held-out test (100 questions): fill precision0.970 (64 of 66)docs/evals/security-questionnaire.md, synthetic library, labels written before any run
    • Held-out test: coverage of answerable questions0.984 (63 of 64)docs/evals/security-questionnaire.md
    • Held-out test: correct abstention when nothing approved fits0.969 (31 of 32)docs/evals/security-questionnaire.md
    • Held-out test: stale approved answers caught4 of 4docs/evals/security-questionnaire.md
    • Invented claims in filled answers (dev + test, 93 fills)0; lexical check 0 missesdocs/evals/security-questionnaire.md
    • Baseline, BM25 top-1 with a threshold from dev: precision / coverage / abstention0.606 / 0.672 / 0.625docs/evals/security-questionnaire.md
    • With the evidence retrieval block (27 Sep, same held-out test, run once): precision / coverage / abstention / stale0.984 (63 of 64) / 0.984 (63 of 64) / 1.000 (32 of 32) / 4 of 4docs/evals/retrieval.md; runs in docs/evals/security-questionnaire/retrieval/
    • Prompt tokens for the 100 held-out questions, BM25 candidates vs the retrieval block218,203 vs 151,150 (-31%)docs/evals/retrieval.md
    • Retrieval alone (reranker top-1, threshold from dev, no model call): precision / coverage / abstention0.864 / 0.797 / 0.844 (BM25: 0.606 / 0.672 / 0.625)docs/evals/retrieval.md
    Latency
    measured: a couple of minutes for the sample sheet through the shared gateway; answers stream in as they finish.
    Verification
    Proof: strongEvery selection and judgement is a separate gateway call with a gateway-signed receipt; the signed record lists them all.
  • Needs more compute

    Wanted: the best setup

    the largest open models, long context

    DeepSeek-V4.1-Flash selects and drafts answers with much more of the evidence library in context, and GLM-5.3-Flash checks each one. Both models support 1M tokens (vendor); the proposed network need serves 131k. Not served yet.

    Models
    • decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks
    • Qwen3.8-27B (NVFP4)
    • DeepSeek-V4.1-Flash
    • GLM-5.3-Flash
    Hardware
    Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.
    Quality evidence
    • answer precision against the BM25 baseline, same protocol as standardnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • Fast option-token selectorGemma-4-26B-A4B-itnot servedA 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway. Hardware: 1x RTX PRO 6000 96 GB (BF16 weights are 49 GB).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks
0 GBProof: partial
Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per questionQwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab)
0.6B + 4B · 12 GBProof: partialSelf-host only
Selector (one call per question) and grounding judge (one call per changed or checked sentence)Qwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Fast selector with option-token probabilities (alternate)Gemma-4-26B-A4B-itgoogle/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
26B (4B active) · 49 GBNo proof yetSelf-host only
Selector and drafter with long contextDeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab)
552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only
Checker, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • Synthetic eval set (docs/evals/security-questionnaire)Apache-2.0

    A fictional vendor's library (50 approved answers, 9 policy and report documents, 2 answers deliberately out of date) and 140 hand-labelled buyer questions: 40 dev, 100 held-out test, written before any model run.

  • Grounding check (tool 17)Apache-2.0

    The per-sentence judge used for both checks: edited answers against the chosen text, and approved answers against the policies they cite.

  • POST /questionnaire/verifyApache-2.0

    Checks a review record's signature against this server's key and, if you send them, the answer hashes. The console also checks the signature in your browser with WebCrypto.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /questionnaire/info, /questionnaire/samples; POST /questionnaire/parse, /questionnaire/answer (SSE or JSON), /questionnaire/export (XLSX or CSV), /questionnaire/verify. Keeps no text.

  • vLLM (selector and judge):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; each call is about 2,200 prompt tokens. Estimate: not run here for this tool.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran on this card, shared with other services.

Latency per lane

  • 30-question sheet, hosted gateway route, 6 questions in flight108.6 s

    Measuredmeasured on our server 2026-09-25: the Corvid Freight demo sheet, 61 receipted calls, answers stream in as each question finishes

  • 100-question eval sheet, gateway route, 4 in flight548.0 s

    Measuredmeasured on our server 2026-09-25: 177 calls, about 2,200 prompt and 120 generated tokens per question (GPU shared with other work)

Notes

  • Select, don't generate. Of 66 answers filled on the held-out test, 22 were the approved text word for word, 29 were whole sentences of it (checked in code, no model), and 15 were edits the grounding judge found supported. The judge rejected no edit in 93 real fills; the revert path is covered by unit tests.
  • Nothing approved, nothing drafted: 31 of 32 test questions with no approved answer went to a person. A top-1 retrieval baseline fills a third of them with the nearest answer.
  • Stale answers: approved answers are checked against the policies they cite. Both out-of-date answers in the demo library were flagged every time they were chosen (4 of 4 on test).
  • The eval is synthetic, written by one author, on one small library, so read it as a check of the mechanism rather than accuracy on real questionnaires.
  • Formats: XLSX (first sheet with a Question header; ID and domain columns detected), CSV and plain text in; XLSX (no formulas) and CSV (formula-like cells neutralised) out. PDFs and Word files are not read: paste their text.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

security-questionnaire/assemble-prompt.md127 lines
# Assemble the Decosa security questionnaire answerer on this machine

You are setting up a questionnaire answerer: it takes our approved answer library (past questionnaire answers, the
trust-center FAQ, our policies, the SOC 2 summary) and a buyer's security questionnaire, and fills each question with
approved text only. An open model chooses the approved answers or policy passages that fit, or says nothing fits; any
sentence it changes is checked against the source by a grounding judge, and approved answers are checked against the
policies they cite. Questions with nothing approved behind them go to a person with nothing drafted. Work step by step,
show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/security-questionnaire.zip (7 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py security-questionnaire` (the api image carries the same bundle under /app/rehearsal/security-questionnaire/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py security-questionnaire --bundle security-questionnaire.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the CSV questionnaire parses into four questions", "the security-policy question (GOV-01) fills from an approved answer", "the encryption question (CEK-01) fills from approved answers naming AES-256 and TLS 1.2"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), used as both the selector and the grounding judge. The answerer is decosa-api
  (Apache-2.0); it reads and writes XLSX and CSV with the Python standard library and needs no GPU of its own.
- Our policies and answers stay on this machine. Bind every port to 127.0.0.1. The service keeps no text: nothing is
  written to disk and logs carry counts only. Keep it that way; do not add request logging.
- Be honest about what it does: it fills only what the library already says, and its checks are model judgements. A
  person reads the sheet before it goes to the buyer. It does not catch omissions (a trimmed answer that drops a
  qualifier), and it does not read PDFs or Word files: export their text first.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus KV cache; an RTX PRO
   6000 96 GB or an RTX 5090 32 GB). Driver 570 or newer. Blackwell cards run NVFP4; on older cards use the FP8 weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for the weights and images.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that contains
  `decosa_api/verticals/questionnaire/` (`main` until one does: v0.1.0 predates it) and `decosa_api/verticals/grounding/`, and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0`; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8` without NVFP4 support).

## 3. docker-compose.yml
Write this in `~/decosa/questionnaire/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_QUESTIONNAIRE_WORKERS: "6"          # questions in flight per run
      DECOSA_QUESTIONNAIRE_MAX_QUESTIONS: "200"  # per run with an API key
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/questionnaire/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.

On the first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up
and never print it. Every model call on the direct route gets a receipt signed with that key (status `attested`), and
each run's review record is signed with it: an attestation by us, the operator, not a proof of computation.

## 4. Smoke test
1. `curl -s localhost:8445/questionnaire/info | head -c 600` lists the statuses, limits and prompt hashes.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"security-questionnaire"}' | jq -r .token)`.
   For more than 60 questions a run, mint an API key with the admin secret instead (`POST /v1/keys`).
3. `curl -s localhost:8445/questionnaire/samples | jq '.samples[0] | {questionnaire, library}' > sample.json`
   (the fictional vendor Tallyloom and a 30-question buyer sheet), then
   `curl -s -N -XPOST localhost:8445/questionnaire/answer -H "authorization: Bearer $T" -H 'accept: text/event-stream' -H 'content-type: application/json' -d @sample.json > run.sse`.
   Expect a `ready` event, then `receipt` and `item` events as questions finish, then `record`, `budget`, `done`.
   On our RTX PRO 6000 the 30 questions took 20-110 s (20 s on a quiet card on 25 Sep 2026) with about 61 model
   calls, and ended with 19-21 `approved`, 0-2 `policy`, 2 `review` (TVM-01 and LOG-01: the library's pen-test and log-retention answers conflict with the
   current policy) and 7 `none` (FedRAMP, bug bounty, PAM, DLP, phishing simulations, FIPS, and the subprocessor question, which
   the library only answers in terms of vendors). Tell me
   what you get; small differences between runs are normal.
4. Export: collect the item events (`grep '"type": "item"' run.sse | sed 's/^data: //' | jq -s .`) and POST
   `{"items": [...], "format": "xlsx", "record": <the record>}` to `/questionnaire/export`; open the file and check
   the Answers sheet has a row per question and blank answers on the `none` rows.
5. `POST /questionnaire/verify` with `{"record": ...}` must show `valid_signature: true` and
   `signed_by_this_server: true`. Change one status in the record and verify again: it must fail.

## 5. Use our own library
- Answers: a CSV or XLSX with columns Question, Answer, and optionally ID, Yes/No, Source (the policy ids the answer
  rests on, separated by `;`) and Origin. Policies and reports: plain text, one document each, with an `id` that
  matches the Source column. Send them as `library: {files: [{name, b64}], documents: [{id, title, kind, text}]}`, or
  preview first with `POST /questionnaire/parse`, which lists citations that point at missing documents.
- The buyer's sheet: the original XLSX as `questionnaire: {file: {name, b64}}`. The answerer uses the first sheet with
  a header containing "question" and picks up ID and domain columns.
- Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local` to use the console against this box.
  Contract: `API_CONTRACT.md`, section "Security questionnaire answerer".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds our policies. If I
ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Fills a buyer's security questionnaire from answers you have already approved, checks every edit against its source, and sends the rest to a person.
Who it's for
Teams in compliance and trust.
Where it runs
Self-host (your policies stay on your machine) or hosted
Key numbers
  • 0.970 (64 of 66) Fill precision (test split, n = 66)
  • 0.984 (63 of 64) Coverage (answerable questions filled correctly) (test split, n = 64)
  • 0.969 (31 of 32) Abstention (unanswerable questions sent to a person) (test split, n = 32)
  • 63.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Self-host (your policies stay on your machine) or hosted
Checks
Receipt per selection and check; signed review record
Output
Structured data · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Does the Decosa security questionnaire answerer write new answers with AI?

No. Decosa's security questionnaire answerer only chooses from answers you have already approved and approved policy passages, and may trim or join them but not add to them. Any sentence it changes goes to a grounding judge; if anything is added, the answer reverts to the approved text word for word. Questions with nothing approved behind them go to a person with no draft.

How accurate is the security questionnaire automation?

On a held-out test of 100 questions against a fictional vendor's library, Decosa's security questionnaire answerer had fill precision of 0.970 (64 of 66), coverage of 0.984 (63 of 64) and abstention of 0.969 (31 of 32), against a BM25 baseline of 0.606, 0.672 and 0.625. The set is synthetic and single-author, and its wording follows the library more closely than real buyer sheets, so expect lower coverage there.

How does it catch stale security questionnaire answers?

Decosa's security questionnaire answerer checks approved answers against the current policies they cite, so an answer saying pen tests run once a year is flagged when the policy now says twice a year. It caught 4 of 4 stale approved answers in the test set, a very small sample. The Lite tier turns this source check off, and then stale approved answers are filled instead of flagged.

What questionnaire formats does it accept?

Decosa's security questionnaire answerer reads the buyer's sheet as XLSX (first sheet with a Question header), CSV, pasted text or JSON, and a library of approved answers (CSV, XLSX or JSON) plus policy documents as text, such as past questionnaires, a trust-center FAQ and a SOC 2 summary. It exports a new XLSX or CSV, not the buyer's own file filled in place. PDFs and Word files are not read.

Do our security policies leave our network?

Not when self-hosted: on the direct route nothing leaves the box, and Decosa's security questionnaire answerer runs on one GPU with Apache-2.0 open weights. The hosted API keeps nothing: the library and questionnaire live in memory for the request, the signed review record holds hashes, source ids and statuses, never text, and logs carry counts. For real policies, self-host.

Should a person still review the filled questionnaire?

Yes. Answers to a security questionnaire can become representations in a contract. Decosa's checks are model judgements and can be wrong: the grounding judge scores 0.59 precision / 0.56 recall on RAGTruth, and a trimmed answer can drop a qualifying sentence without being caught. Every answer carries its source and receipts so a reviewer can trace it before the sheet is sent.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Security questionnaire answerer

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.