Map evidence to NIST 800-171
A gap map of all 320 NIST 800-171 objectives, each marked present, partial or missing with quotes from your evidence, plus a draft POA&M.
Built on: Typed judgment, Grounding, Signed record, Agent flight recorder
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the catalog, artefact checks, POA&M rules and signed record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the cmmc / nist 800-171 evidence map API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- cmmc-evidence-map
Use the hosted API
# Decosa CMMC / NIST 800-171 evidence map: use the hosted API (synthetic sets only)
You are wiring Decosa's CMMC evidence map into this project. It takes a contractor's SSP statements and evidence
artefacts (policies, config exports, logs, records, scans, and on the hosted API only the bundled sample screenshot and
test run) and, for each NIST SP 800-171 Rev 2 requirement and each NIST SP 800-171A objective, says evidence present,
partial or missing, with the lines quoted and the artefact checks (draft, stale, undated, out of scope). It returns a
Markdown gap map, a draft POA&M with the 32 CFR 170.21 eligibility of each row, an evidence index with every artefact's
SHA-256 and a record signed by the server, with a signed receipt for every model call. Use only what is listed below.
If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic sets only.** An SSP and its evidence are often CUI, which must not go to a hosted API.
Every request must carry `"synthetic": true` (HTTP 422 otherwise); text with CUI banner or portion markings, a CUI
designation block or a distribution statement B-F is refused with 422 before any model call; screenshots are accepted
only when they are the bundled sample images. Real SSPs belong on a self-hosted box (see the self-host prompt).
- It is an evidence-gap map, not an assessment, a certification or a score. Never show its output as "MET", "compliant"
or as an SPRS score. Fields such as `certify`, `score` or `affirm` in a request get HTTP 400. Show gaps first.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in
code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "cmmc-evidence-map"}` returns `{"token", "expires_at", "budget"}`.
Sessions per IP are limited; over a limit you get HTTP 429 with `Retry-After`.
3. A map needs about 700 generated tokens of budget per requirement plus about 370 per objective and 700 per screenshot (402 otherwise; the budget check uses these ceilings, actual use is lower). One
map at a time per demo token (409 while one is going). Up to 12 requirements per run.
## Endpoints
- `POST /cmmc/map` (token). Body:
```json
{"system": {"name": "Example CUI enclave", "assessment_date": "2026-09-15", "scope": ["ex-fs01", "ex-wks*"]},
"ssp": {"statements": [{"requirement": "3.1.8", "status": "implemented", "text": "...", "artefacts": ["A1"]}]},
"artefacts": [{"id": "A1", "title": "Account lockout settings export", "kind": "config", "date": "2026-08-30",
"system": "ex-fs01", "status": "final", "text": "..."}],
"requirements": ["3.1.8"], "synthetic": true, "stream": false}
```
`kind`: policy, procedure, config, log, screenshot, scan, record, inventory, test_run, other. `ssp` may instead be
`{"text": "..."}` (the SSP pasted as text, split on requirement ids such as `3.1.8` or `AC.L2-3.1.8`). Up to 30
artefacts (20,000 characters each, 120,000 in total). `{"sample": "harbor-precision"}` runs a built-in sample.
- JSON answer: `{status: gaps | no_gaps_in_evidence, summary, counts, requirements: [{id, cmmc_id, points, poam:
{eligible, rule, why}, ssp, notes, status: present | partial | missing | not_applicable, objectives: [{id, text, status,
reason, issues, quotes: [{line, artefact, text, transcribed, issues}], typed, grounding}]}], poam: {note, rows},
evidence_index: [{id, sha256, issues, certificate?, cited_for, counted_for}], estimate: {shown: false, why}, receipts,
gap_map_md, record}`.
- SSE: send `Accept: text/event-stream` (or `"stream": true`): `ready`, `receipt`, `artefacts`, `requirement` (one per
requirement), `report`, `budget`, `done`.
- `POST /cmmc/check` (no token, no model): the same body → the artefact checks and the CUI guard only.
- `GET /cmmc/catalog?requirement=3.10.3` or `?family=PE` (no token): requirements, objectives, points, POA&M rules.
- `POST /record/verify` (no token): the `record` → `{ok, checks, summary}`.
- `GET /cmmc/info`, `GET /cmmc/samples` (no token): the rules, the catalog's sources, limits and three samples.
## Errors
400 bad input (the message names the field or limit, or refuses a certify or score field), 401/403 token, 402 budget,
409 a map already going on this demo token, 413 body over 8 MB, 422 not synthetic, marked CUI or a screenshot that is
not a sample, 429 busy (`Retry-After`).
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa CMMC / NIST 800-171 evidence map: run it yourself (containers)
You are setting up the Decosa CMMC evidence map on this machine, so the SSP and its evidence (often CUI) never leave it.
For each NIST SP 800-171 Rev 2 requirement and NIST SP 800-171A objective it says evidence present, partial or missing,
with the lines quoted from the artefacts, flags artefacts that cannot count, drafts a POA&M with the 32 CFR 170.21
eligibility of each row, and signs an evidence index with every artefact's SHA-256. Nothing is sent to Decosa's hosted
API. It is an evidence-gap map, not an assessment or a score: an assessor decides MET, the senior official affirms.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/cmmc-evidence-map.zip (66 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py cmmc-evidence-map` (the api image carries the same bundle under /app/rehearsal/cmmc-evidence-map/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py cmmc-evidence-map --bundle cmmc-evidence-map.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "encryption is only planned in the SSP, so 3.13.11 has no evidence", "3.1.8 has a planted gap, so it is not fully evidenced", "3.3.1 has a planted gap, so it is not fully evidenced"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CMMC_SYNTHETIC_ONLY=0` (so this box accepts real SSPs), `DECOSA_CMMC_MAX_REQUIREMENTS=110`, and bind every
port to 127.0.0.1. In real-data mode the api refuses to map (503) if model calls would go to a hosted endpoint; never
set the gateway route on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/cmmc/info` shows `data_rule.synthetic_only: false`,
`data_rule.real_data_ready: true` and 320 objectives; `GET /attest/signing-key` shows this box's public key. Show me
the key: it is what an assessor pins to verify my evidence index.
5. Smoke test: get a token with `POST /demo/session {"vertical":"cmmc-evidence-map"}` and send
`{"sample": "harbor-precision", "synthetic": true}` to `POST /cmmc/map`. Expect 3.13.11 missing (planned in the SSP),
none of 3.1.8, 3.3.1, 3.5.3, 3.5.7, 3.10.3 or 3.14.2 fully evidenced, artefacts A4 stale, A5 draft and A9 out of
scope, the test run T1 verified, the screenshot S1 transcribed, the POA&M row for 3.10.3 marked "never on a POA&M"
and no estimate. Send the record to `POST /record/verify`: `ok` must be true.
6. Report back: the public key, the counts, the gaps with their reasons, and how long it took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds CUI.
If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/cmmc-evidence-map-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMlite tierRuns with a smaller tier
The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"cmmc-evidence-map"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py cmmc-evidence-map
Download the mock-data bundle (66 KB, 21 checks)expected.json
A fictional 38-person machine shop preparing for a CMMC Level 2 self-assessment. The SSP claims every requirement is implemented except 3.13.11 (planned). Planted: a lockout setting that never locks (3.1.8), a year-old audit log export (3.3.1), MFA shown only in a screenshot where the all-users policy is report-only (3.5.3), a draft password standard (3.5.7), a visitor log with no escort column (3.10.3, which can never go on a POA&M), encryption planned (3.13.11) and an anti-malware export from a system outside the scope (3.14.2). 3.1.1 and 3.11.2 are fully evidenced, 3.1.1 with a verified Decosa test run. The run must never call a planted gap fully evidenced, flag the stale, draft and out-of-scope artefacts, verify the test run, mark 3.10.3 as never on a POA&M, show no score, and sign a record that verifies and fails when changed. (The hosted service's refusals of unmarked and CUI-marked sets are covered by the smoke test and the unit tests, since a self-hosted box in real-data mode accepts both.)
What the rehearsal checks
- encryption is only planned in the SSP, so 3.13.11 has no evidence
- 3.1.8 has a planted gap, so it is not fully evidenced
- 3.3.1 has a planted gap, so it is not fully evidenced
- 3.5.3 has a planted gap, so it is not fully evidenced
- 3.5.7 has a planted gap, so it is not fully evidenced
- 3.10.3 has a planted gap, so it is not fully evidenced
- 3.13.11 has a planted gap, so it is not fully evidenced
- 3.14.2 has a planted gap, so it is not fully evidenced
- 3.10.3 (escort visitors) can never go on a POA&M
- 3.13.11 may go on a POA&M only when encryption is employed but not FIPS-validated
- the two-year-old audit log export is flagged stale
- the password standard is flagged as a draft
- the anti-malware export from a system outside the scope is flagged
- the Decosa test-run certificate verifies
- the screenshot was read by the vision model
- some objectives have evidence present
- no score or estimate is shown for 9 of 110 requirements
- the gap map lists the draft POA&M
- the signed record verifies
- a record with its status changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic: the company, people, systems and evidence are invented (decosa_api/verticals/cmmc/synth.py and pool.py); the screenshot is drawn by scripts/cmmc_make_sample_png.py; the test-run certificate is a real Decosa run against saucedemo.com, Sauce Labs' public test site. NIST SP 800-171 and 800-171A are US government works. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa CMMC / NIST 800-171 evidence map: run it yourself (containers)
You are setting up the Decosa CMMC evidence map on this machine, so the SSP and its evidence (often CUI) never leave it.
For each NIST SP 800-171 Rev 2 requirement and NIST SP 800-171A objective it says evidence present, partial or missing,
with the lines quoted from the artefacts, flags artefacts that cannot count, drafts a POA&M with the 32 CFR 170.21
eligibility of each row, and signs an evidence index with every artefact's SHA-256. Nothing is sent to Decosa's hosted
API. It is an evidence-gap map, not an assessment or a score: an assessor decides MET, the senior official affirms.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/cmmc-evidence-map.zip (66 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py cmmc-evidence-map` (the api image carries the same bundle under /app/rehearsal/cmmc-evidence-map/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py cmmc-evidence-map --bundle cmmc-evidence-map.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "encryption is only planned in the SSP, so 3.13.11 has no evidence", "3.1.8 has a planted gap, so it is not fully evidenced", "3.3.1 has a planted gap, so it is not fully evidenced"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CMMC_SYNTHETIC_ONLY=0` (so this box accepts real SSPs), `DECOSA_CMMC_MAX_REQUIREMENTS=110`, and bind every
port to 127.0.0.1. In real-data mode the api refuses to map (503) if model calls would go to a hosted endpoint; never
set the gateway route on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/cmmc/info` shows `data_rule.synthetic_only: false`,
`data_rule.real_data_ready: true` and 320 objectives; `GET /attest/signing-key` shows this box's public key. Show me
the key: it is what an assessor pins to verify my evidence index.
5. Smoke test: get a token with `POST /demo/session {"vertical":"cmmc-evidence-map"}` and send
`{"sample": "harbor-precision", "synthetic": true}` to `POST /cmmc/map`. Expect 3.13.11 missing (planned in the SSP),
none of 3.1.8, 3.3.1, 3.5.3, 3.5.7, 3.10.3 or 3.14.2 fully evidenced, artefacts A4 stale, A5 draft and A9 out of
scope, the test run T1 verified, the screenshot S1 transcribed, the POA&M row for 3.10.3 marked "never on a POA&M"
and no estimate. Send the record to `POST /record/verify`: `ok` must be true.
6. Report back: the public key, the counts, the gaps with their reasons, and how long it took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds CUI.
If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/cmmc-evidence-map-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsCMMC / NIST 800-171 evidence map on GeForce RTX 5090: use the Standard · one GPU for the model (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case.
Standard · one GPU for the model (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Catalog, point values and POA&M rules, artefa...: decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Screenshot transcription, lines per objective...: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for CMMC / NIST 800-171 evidence map, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up CMMC / NIST 800-171 evidence map on my hardware Fetch https://decosa.ai/prompts/cmmc-evidence-map-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=cmmc-evidence-map) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one GPU for the model (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Catalog, point values and POA&M rules, artefa...: decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27), CPU - Screenshot transcription, lines per objective...: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/cmmc-evidence-map-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa CMMC / NIST 800-171 evidence map: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa CMMC / NIST 800-171 evidence map on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/cmmc-evidence-map.zip (66 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py cmmc-evidence-map` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "encryption is only planned in the SSP, so 3.13.11 has no evidence", "3.1.8 has a planted gap, so it is not fully evidenced", "3.3.1 has a planted gap, so it is not fully evidenced"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Screenshot transcription, lines per objective, the typed judgment per objective, and the grounding judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py cmmc-evidence-map`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Built from the same parts as the measured tools (Qwen3.8-27B and CPU code); not run on the Mac yet. Screenshots need the model's vision input, which we have not checked on the Mac build. - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 150 s · ~$0.008 per run · 33 receipts
Loading the nightly status…
Self-host: verified 26 Sep 2026 · fresh clone, compose up, sample against local model servers
Measured cost to run: about $0.26 per 100 requirements (hosted, 26 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_CMMC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 21 of 21 checks in 19.8 s, the smoke module passed in 5.3 s with 18 attested receipts, unmarked and CUI-marked sets were accepted as real-data mode should, and the 10 held-out test companies scored 235 of 286 with no false present. Model-server startup itself not re-verified.
Known limits (6)
- Synthetic only on the hosted demo, and every number here comes from our own synthetic companies over 14 of the 110 requirements; not run on real SSPs or against an assessor's labels.
- Reads text, config exports and single screenshots; the hosted demo reads only the bundled sample screenshot. Scanned PDFs and Word binders are not read.
- It errs down: recall of 'present' is 72% on the test set, so some evidenced objectives come back partial.
- 'Missing' means not in what was sent. It does not interview, test or examine systems as an assessor does.
- No SPRS score. A self-assessed estimate appears only when all 110 requirements are mapped in one run (self-hosted), labelled as an estimate.
- NIST SP 800-171 Rev 2 and CMMC Level 2 only; no Level 1, Level 3 or Rev 3.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the catalog, artefact checks, POA&M rules and signed record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the cmmc / nist 800-171 evidence map API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Each NIST SP 800-171 objective checked against the evidence you send: present, partial or missing, with quotes. A gap map, never a score.
For small defence contractors preparing for CMMC Level 2, and the consultants (RPOs, MSPs) who prepare them. Send the SSP's statements and the evidence behind them: policies, config exports, logs, records, scans, a screenshot, a verified Decosa test run. Code loads the 110 NIST SP 800-171 Rev 2 requirements and 320 NIST SP 800-171A objectives from NIST's own files, with the 32 CFR 170.24 point values and the 170.21 POA&M rules, and checks every artefact: final or draft, dated and fresh, from a system in the SSP's scope, and, for test runs, the certificate's signature and assertions. Qwen3.8-27B transcribes screenshots, lists the lines that bear on each objective, gives a typed judgment (present, partial, contradicted, missing), and the grounding judge checks that the quotes carry a 'present'. Rules in code decide, and every doubt goes down: the SSP counts only for objectives about something being defined or identified, and a draft, stale or out-of-scope artefact makes an objective partial at most. Out come a gap map (gaps first), a draft POA&M that marks what may never go on one, and a signed evidence index with each artefact's SHA-256 that an assessor can check. It never says MET and never states an SPRS score.
- Deployment
- Self-host first
- Regulatory
- Checked 26 Sep 2026 against primary sources (links under Tools). 32 CFR Part 170 (the CMMC Program rule, 89 FR 83092, effective 16 Dec 2024) sets Level 2 as the 110 requirements of NIST SP 800-171 Rev 2, assessed per NIST SP 800-171A (June 2018), both incorporated by reference in 170.2; NIST's Rev 3 (May 2024) is not what CMMC assesses. 170.24 defines MET, NOT MET and N/A, requires evidence in final form (no drafts or unapproved policies), and sets the point values (5, 3, 1; 3.5.3 and 3.13.11 variable). 170.21 limits a POA&M: a score of at least 0.8 of 110, only 1-point requirements (3.13.11 at 3 points when encryption is employed but not FIPS-validated), never 3.1.20, 3.1.22, 3.10.3, 3.10.4, 3.10.5 or 3.12.4, closed out within 180 days. The DFARS rule (DFARS Case 2019-D041, 90 FR 43560) put CMMC into DoD contracts from 10 Nov 2025 (Phase 1). A pause of the move to Phase 2 (Level 2 C3PAO certification in new contracts from 10 Nov 2026) for a 60-day review from 13 Jul 2026 (DoD memo 26-P-1023) is reported by secondary sources; we did not find the memo itself, so that is unverified here, and we do not know the review's outcome. NIST's 800-171A CSV labels 3.10.1 as Personnel Security; the catalog takes the family from the section number. An SSP and its evidence are often CUI: self-host first; the hosted demo takes synthetic sets only and refuses marked CUI. This is not legal advice, an assessment or a certification: an assessor decides MET, and the senior official affirms in SPRS. Model licence: Apache-2.0 (Qwen3.8-27B). NIST publications and 32 CFR are US government works.
Text description
A contractor's SSP statements and evidence artefacts go to the evidence map. In code, each artefact is checked: final not draft, dated and fresh, from a system in the SSP's scope, and Decosa test-run certificates are verified; on the hosted route, text with CUI markings is refused first. Qwen3.8-27B transcribes screenshots, lists the lines that bear on each NIST SP 800-171A objective, gives a typed judgment, and the grounding judge checks that the quotes carry a present. Rules in code give evidence present, partial or missing, gaps first, and never count a doubt as present. Outputs: a gap map, a draft POA&M with the 32 CFR 170.21 eligibility of every row, and a signed evidence index with each artefact's hash and no artefact text. On the hosted route, which takes synthetic sets only, every model call gets a gateway-signed receipt. Self-hosted, everything stays on your machine.
At a glance
- Data retention
- Nothing stored: the SSP and artefacts live in memory for the request. The signed index holds hashes, statuses and receipt ids, never artefact text; logs carry counts only.
- What leaves the box
- Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and only synthetic sets are accepted (marked CUI is refused before any call). Self-hosted on the direct route: nothing leaves the box, and in real-data mode the service refuses to run if model calls would.
- What it will not do
- Say MET, certify, score or affirm. It maps evidence to objectives, lists the gaps first, and leaves MET to the assessor and the affirmation to the senior official.
- POA&M rules
- Every draft row carries its 32 CFR 170.21 eligibility: 1-point requirements only (3.13.11 conditionally), never 3.1.20, 3.1.22, 3.10.3, 3.10.4, 3.10.5 or 3.12.4.
- Input formats
- JSON: the system in scope (name, assessment date, scope list), SSP statements (or the SSP pasted as text), and up to 30 artefacts of up to 20,000 characters (120,000 in total), plus screenshots and Decosa test-run certificates. Up to 12 requirements per hosted run, 110 self-hosted.
- Typical run
- The clean sample: a few dozen model calls and a fraction of a cent at the gateway list price, from under a minute to a few minutes on the shared gateway; self-hosted, seconds. Each run shows its own measured cost.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Lite
catalog and artefact checks only, no GPU
GET /cmmc/catalog gives every requirement's objectives, points and POA&M rule; POST /cmmc/check flags draft, stale, undated and out-of-scope artefacts and verifies test-run certificates. No reading of the evidence, so no statuses.
- Models
- decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)
- Hardware
- Any CPU
- Quality evidence
- Catalog against NIST's CSVs and 32 CFR 170110 requirements, 320 objectives; 42 five-point, 14 three-point, 2 variable, 52 one-point; the six never-POA&M requirements; unit-testedtests/test_cmmc.py, 26 Sep 2026
- Artefact-check accuracy on real evidencenot measured yetnot measured yet
- Latency
- no model call; milliseconds on CPU (not separately timed)
- Verification
- No proof yetNo model call, so no receipts; the checks are deterministic code.
- In the hosted demo
Standard
one GPU for the model (hosted demo)
Qwen3.8-27B reads the evidence per requirement and objective (lines, typed judgment, grounding) and transcribes screenshots; the catalog, artefact checks, status rules, POA&M rules and signed index are code. This is what the hosted demo runs, on synthetic sets only.
- Models
- decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Objective status, 3 classes (10 held-out synthetic companies, 286 objectives)237 / 286docs/evals/cmmc-evidence-map.md, test split, 26 Sep 2026
- Partial or missing objectives called present (false present)0 / 121docs/evals/cmmc-evidence-map.md, test split
- Requirements called fully evidenced that are not0 / 59docs/evals/cmmc-evidence-map.md, test split
- Present objectives called present (recall)119 / 165docs/evals/cmmc-evidence-map.md, test split
- Quotes verbatim from the input446 / 446docs/evals/cmmc-evidence-map.md, test split
- Objective status, self-hosted direct route (same 286 test objectives)235 / 286, false present 0 / 121docs/evals/cmmc-evidence-map.md, self-hosted test run
- Latency
- measured: under a minute per company self-hosted on the direct route; minutes per company under a heavily shared gateway, and from under a minute to several minutes for the demo and smoke samples.
- Verification
- Proof: strongEvery model call is a separate gateway call with a gateway-signed receipt; the signed index lists them all.
Also runs on
- Evidence binders as PDFs (document reader)A licence-clean document OCR and layout model (not chosen)not builtRead SSPs and evidence binders as PDF or Word with page references in quotes. Waits on the document reader block (page 48); not built. Hardware: 1x RTX PRO 6000 96 GB (estimate).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU)decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27) 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
Screenshot transcription, lines per objective, the typed judgment per objective, and the grounding judgeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | Standard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Scanned PDF and evidence-binder reader (wanted)A licence-clean document OCR and layout model (not chosen) No proof yetSelf-host only | Alternate | n/a | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- 32 CFR Part 170, CMMC Program (final rule, 89 FR 83092, 15 Oct 2024) (opens in a new tab)US federal regulation (public domain)
170.24 scoring (MET, NOT MET, N/A; final-form evidence; point values) and 170.21 POA&M rules, read on eCFR 26 Sep 2026.
- DFARS Case 2019-D041 (final rule, 90 FR 43560, 10 Sep 2025) (opens in a new tab)US federal regulation (public domain)
CMMC in DoD contracts from 10 Nov 2025 (Phase 1).
- NIST SP 800-171A assessment procedures (CSV) (opens in a new tab)US government work (public domain)
The 320 assessment objectives. Rebuilt by scripts/cmmc_build_catalog.py; CSV SHA-256 in GET /cmmc/info.
- NIST SP 800-171 Rev 2 security requirements (CSV) (opens in a new tab)US government work (public domain)
The 110 requirements and their basic or derived designation.
The reported 13 Jul 2026 pause of Phase 2 (memo 26-P-1023, 60-day review). Unverified against a primary source.
- Decosa test runs (27)AGPL-3.0-or-later
Signed certificates of an end-to-end test (here: a sign-in page refusing a locked-out account and an unauthenticated visit) as evidence that a control works; verified before they are read.
- scripts/eval_cmmc.py and docs/evals/cmmc-evidence-map.mdApache-2.0
The synthetic company generator, the dev and test runs and every result row.
- GET /cmmc/catalog, POST /cmmc/check and POST /record/verifyApache-2.0
Look up requirements, objectives, points and POA&M rules, and run the artefact checks, with no model call; verify a signed evidence index anywhere.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /cmmc/info, /cmmc/catalog, /cmmc/samples; POST /cmmc/map (SSE or JSON), /cmmc/check; POST /record/verify. Keeps no artefact text.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo, the eval and the smoke test ran through the shared gateway on this card.
- 1x RTX 5090 32 GB Fits
Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool.
- CPU only Fits
The lite tier (GET /cmmc/catalog, POST /cmmc/check: catalog, POA&M rules, artefact checks) needs no GPU.
Latency per lane
- one company, 8 requirements, hosted gateway route236.7 s
Measuredmeasured on our server 2026-09-26: mean over the 10 test companies, 3 in parallel, gateway shared with other workloads
- harbor-precision sample, 9 requirements with a screenshot, hosted gateway route354.7 s
Measuredmeasured on our server 2026-09-26 while recording the Watch run under heavy gateway load; an earlier run of the same sample took 25.8 s
- bluefin-clean sample, 3 requirements (the smoke test), hosted gateway route149.5 s
Measuredmeasured on our server 2026-09-26 by scripts/smoke/run_all.py under load; 16.2 s in an earlier, quieter run
- one company, 8 requirements, self-hosted direct route24.1 s
Measuredmeasured on our server 2026-09-26 in the self-host check: mean over the 10 test companies, 3 in parallel, typed judgment by log-probabilities
- catalog and artefact checks only (GET /cmmc/catalog, POST /cmmc/check)n/a
Measuredno model call; milliseconds on CPU (not separately timed)
Notes
- On 10 held-out synthetic companies (80 requirements, 286 objectives, worded differently from the dev set) it gave the labelled status for 237 objectives. It called none of the 121 partial or missing objectives present, and no requirement fully evidenced that was not (0 of 59).
- It errs down: 47 of its 49 test errors call an objective less evidenced than its label (recall of 'present' is 119 of 165). A preparer re-checks some objectives that were in fact evidenced.
- Every quote is a line of an artefact sent (446 of 446 on test), picked by id; 204 of 209 present or partial calls quote a planted evidence line.
- Draft, stale, undated and out-of-scope artefacts are caught in code, not by the model, and a test-run certificate that fails verification is not read at all.
- Everything is measured on synthetic companies written by the agent that built the checker, over 14 of the 110 requirements. It has not been run on real SSPs or against an assessor's labels.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa CMMC / NIST 800-171 evidence map on this machine
You are setting up an evidence-gap map for a small defence contractor (or the consultant preparing one) ahead of a CMMC
Level 2 self-assessment or certification assessment. It takes the SSP's statements and the evidence artefacts (policies,
config exports, logs, records, scans, screenshots, Decosa test-run certificates) and, for each NIST SP 800-171 Rev 2
requirement and each NIST SP 800-171A objective, says evidence present, partial or missing, with the lines quoted. It
flags artefacts that cannot count (drafts, stale exports, systems outside the scope), drafts a POA&M with the
32 CFR 170.21 eligibility of every row, and signs an evidence index with each artefact's SHA-256. Work step by step,
show me each command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/cmmc-evidence-map.zip (66 KB, 21 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py cmmc-evidence-map` (the api image carries the same bundle under /app/rehearsal/cmmc-evidence-map/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py cmmc-evidence-map --bundle cmmc-evidence-map.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "encryption is only planned in the SSP, so 3.13.11 has no evidence", "3.1.8 has a planted gap, so it is not fully evidenced", "3.3.1 has a planted gap, so it is not fully evidenced"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0): the lines per objective, the typed judgment, the grounding check, and reading
screenshots. The catalog (NIST SP 800-171 Rev 2 and 800-171A, US government works), the point values and POA&M rules
(32 CFR 170), the artefact checks, the status rule and the signed record are decosa-api (AGPL-3.0-or-later), CPU only.
- An SSP and its evidence are often CUI. They stay on this machine. Bind every port to 127.0.0.1. The service keeps no
artefact text: nothing is written to disk and logs carry counts only. Keep it that way; do not add request logging.
- It is an evidence-gap map, not an assessment, a certification or a score. It never states an SPRS score; it shows a
"self-assessed estimate" only when all 110 requirements are mapped in one run. An assessor decides MET; the senior
official affirms.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; we measured
on an RTX PRO 6000 96 GB; an RTX 5090 32 GB should fit, not run for this tool). Driver 570 or newer. Blackwell
cards run NVFP4; on older cards use `Qwen/Qwen3.8-27B-FP8`. Screenshots need the model's image input: vLLM serves it
for this model; check that `/v1/models` lists it and that the smoke test's screenshot step passes.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that
contains `decosa_api/verticals/cmmc/` (`main` until one does), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
`482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on a card without NVFP4.
- The catalog ships inside the image (`decosa_api/verticals/cmmc/data/nist171.json`). To rebuild it from NIST's own CSVs,
see `scripts/cmmc_build_catalog.py`; `GET /cmmc/info` shows the source URLs and the CSVs' SHA-256.
## 3. docker-compose.yml
Write this in `~/decosa/cmmc/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_CMMC_SYNTHETIC_ONLY: "0"
DECOSA_CMMC_MAX_REQUIREMENTS: "110"
DECOSA_CMMC_WORKERS: "4"
DECOSA_BUDGET_LLM_TOKENS: "400000"
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/cmmc/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
`DECOSA_CMMC_SYNTHETIC_ONLY: "0"` lets this box take real SSPs. In that mode the api refuses to map (HTTP 503) if model
calls would leave the box: on the gateway route, or to a direct endpoint that is not this host, a private address or a
container name like `llm`. If your model endpoint is yours but elsewhere (your own GovCloud tenancy), set
`DECOSA_CMMC_MODEL_SELF_HOSTED: "1"` to say so. Never set `DECOSA_LLM_ROUTE=gateway` on this box.
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a host folder:
the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by root, which stops
the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Start everything:
`docker compose up -d`.
On the first start the api creates this box's Ed25519 key in `decosa-data` (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private; never print it. Every model call on the
direct route gets a receipt signed with that key (status `attested`): an attestation by me, the operator, not a proof
of computation. The public key (`GET /attest/signing-key`) is what an assessor pins to verify my evidence index.
## 4. Smoke test
1. `curl -s localhost:8445/cmmc/info | jq '{d: .data_rule, m: .model.judgment_method, n: .catalog.objectives}'` shows
`synthetic_only: false`, `real_data_ready: true`, `logprobs` and `320`.
2. `curl -s 'localhost:8445/cmmc/catalog?requirement=3.10.3' | jq .poam` shows `eligible: false` (never on a POA&M).
3. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"cmmc-evidence-map"}' | jq -r .token)`.
4. `curl -s -XPOST localhost:8445/cmmc/map -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"harbor-precision","synthetic":true}' > res.json`.
Expect: 3.13.11 `missing` (the SSP says planned), none of 3.1.8, 3.3.1, 3.5.3, 3.5.7, 3.10.3, 3.14.2 fully present,
artefact A4 `stale`, A5 `draft`, A9 `out_of_scope`, T1's certificate `ok: true`, S1 `transcribed: true`, the
POA&M row for 3.10.3 with `poam_eligible: false`, and `estimate.shown: false`. Statuses read by the model can differ
between runs; the artefact checks and POA&M rules cannot.
5. `jq .record res.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @- | jq .ok`
must print `true`. Change one `requirement` entry's `status` to `present` and verify again: it must fail.
6. If decosa-api's source is at hand, `python scripts/rehearse.py cmmc-evidence-map --base-url http://127.0.0.1:8445`
runs this and prints PASS or FAIL per property.
7. Time it and tell me what you measure. The Harbor sample took about 26-70 s through the shared hosted gateway.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /cmmc/map` per family or per
SSP section and keep `gap_map_md`, the POA&M rows and `record` with the assessment folder. `POST /cmmc/check` runs the
artefact checks alone with no model call. Contract: `API_CONTRACT.md`, section "Changes (cmmc-evidence-map, 2026-09-26)".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds CUI. If I ask for it,
follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
9 laws, rules and guidance pages cited; 9 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Each NIST SP 800-171 objective checked against the evidence you send: present, partial or missing, with quotes. A gap map, never a score.
- Who it's for
- Small defence contractors preparing a CMMC Level 2 assessment, and the RPOs and MSPs who prepare them.
- Where it runs
- Self-host for real SSPs (CUI stays on your box); the hosted demo takes synthetic sets only and refuses marked CUI
- Key numbers
- 237 / 286 Objective status, 3 classes (test split, n = 286)
- 0 / 121 Partial or missing objectives called present (false present) (test split, n = 121)
- 0 / 59 Requirements called fully evidenced that are not (test split, n = 59)
- 149.5 s Median end-to-end run, hosted (QA sweep 2026-09-26)
- Models
- Qwen3.8-27B reads the evidence (lines per objective, typed judgment, grounding) and transcribes screenshots; the catalog, point values, POA&M rules and artefact checks are plain code
- Where
- Self-host for real SSPs (CUI stays on your box); the hosted demo takes synthetic sets only and refuses marked CUI
- Checks
- Receipt per model call; every quote is a line of an artefact you sent; test-run certificates verified; signed hash-chained evidence index with no artefact text
- Industry
- Compliance and trust · Public sector
- Runs
- Self-host
- Output
- Signed record or verdict · Structured data
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Typed judgment · Grounding · Signed record · Agent flight recorder
Questions people ask
Does it tell me whether we pass CMMC Level 2?
No. It maps the evidence you send to each NIST SP 800-171A objective and lists the gaps. An assessor decides MET or NOT MET, and your senior official makes the affirmation. It never states an SPRS score; a labelled self-assessed estimate appears only when all 110 requirements are mapped in one run.
Does our SSP or evidence leave our network?
Not when you self-host: the model runs on your GPU, and in real-data mode the service refuses to run if model calls would go to a hosted endpoint. The hosted demo takes synthetic sets only and refuses text with CUI markings before any model call.
How often does it call a gap 'present'?
On 10 held-out synthetic companies it called none of 121 partial or missing objectives present, and got 237 of 286 objectives right. It errs down: some evidenced objectives come back partial. These are synthetic sets, not real SSPs.
Which requirements can go on a POA&M?
Under 32 CFR 170.21, only 1-point requirements (and 3.13.11 when encryption is employed but not FIPS-validated), with a score of at least 88 of 110, and never 3.1.20, 3.1.22, 3.10.3, 3.10.4, 3.10.5 or 3.12.4. Each draft POA&M row says which applies.
Which NIST 800-171 revision does it use?
Rev 2 with the SP 800-171A objectives (June 2018), which is what 32 CFR 170 incorporates for CMMC Level 2. Rev 3 is not covered.
What evidence does it catch that a spreadsheet misses?
Draft or unapproved policies, exports older than a year, artefacts from systems outside the SSP's scope, and SSP claims no artefact shows. Each is flagged in code with the rule behind it, and every quote is a line of the artefact you sent.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about CMMC / NIST 800-171 evidence map
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…