65 · Compliance and trust · Public sector · live
CMMC / NIST 800-171 evidence map
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Objective status, 3 classes237 / 286test splitn = 28610 synthetic companies, 80 requirements; present, partial, missing.
- Partial or missing objectives called present (false present)0 / 121test splitn = 121
- Requirements called fully evidenced that are not0 / 59test splitn = 59
- Present objectives called present (recall)119 / 165test splitn = 165It errs down: 47 of 49 errors call an objective less evidenced than its label.
- Quotes verbatim from the input446 / 446test splitn = 446
- Present or partial calls quoting a planted evidence line204 / 209test splitn = 209
- Requirement roll-up accuracy64 / 80test splitn = 80
- Objective status, self-hosted direct route235 / 286test splitn = 286Log-probabilities instead of votes; false present 0 / 121. One rule change made on the self-hosted dev set first.
- Objective status, 3 classes (dev)238 / 299dev (tuned on)n = 299215 / 299 before the dev changes; false present on dev 3 / 102.
Dataset
Synthetic small defence contractors from the vertical's own generator (invented companies, people and systems): 8 requirements each from a pool of 14 NIST SP 800-171 Rev 2 requirements (51 objectives), each with at most one planted gap (evidence missing, planned, stale, out of scope, draft, shown in part, contradicted, one line dropped). Dev 10 companies (299 objectives, one phrasing), test 10 companies (286 objectives, every evidence line worded differently), run once after the dev work was frozen.
Caveats
- Everything is synthetic, from templates written by the same agent that wrote the prompts and rules: evidence the mechanisms work, not accuracy on real SSPs.
- No RPO or assessor labels; 14 of the 110 requirements are exercised.
- Labels count only a requirement's own artefacts; the model sometimes uses evidence filed under another requirement.
- Screenshot reading is shown on one bundled image, not measured.
- Measured on a gateway shared with other workloads, so the latencies are high and vary.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 150 s
- Receipts
- 33
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.008
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_CMMC_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 21 of 21 checks in 19.8 s, the smoke module passed in 5.3 s with 18 attested receipts, unmarked and CUI-marked sets were accepted as real-data mode should, and the 10 held-out test companies scored 235 of 286 with no false present. Model-server startup itself not re-verified.
Rehearsal bundle: cmmc-evidence-map.zip (66 KB, 21 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Synthetic only on the hosted demo, and every number here comes from our own synthetic companies over 14 of the 110 requirements; not run on real SSPs or against an assessor's labels.
- Reads text, config exports and single screenshots; the hosted demo reads only the bundled sample screenshot. Scanned PDFs and Word binders are not read.
- It errs down: recall of 'present' is 72% on the test set, so some evidenced objectives come back partial.
- 'Missing' means not in what was sent. It does not interview, test or examine systems as an assessor does.
- No SPRS score. A self-assessed estimate appears only when all 110 requirements are mapped in one run (self-hosted), labelled as an estimate.
- NIST SP 800-171 Rev 2 and CMMC Level 2 only; no Level 1, Level 3 or Rev 3.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU)decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)AGPL-3.0-or-later
- Screenshot transcription, lines per objective, the typed judgment per objective, and the grounding judgeQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · catalog and artefact checks only, no GPU (2)
- Catalog against NIST's CSVs and 32 CFR 170: 110 requirements, 320 objectives; 42 five-point, 14 three-point, 2 variable, 52 one-point; the six never-POA&M requirements; unit-testedtests/test_cmmc.py, 26 Sep 2026
- Artefact-check accuracy on real evidence: not measured yetnot measured yet
Standard · one GPU for the model (hosted demo) (6)
- Objective status, 3 classes (10 held-out synthetic companies, 286 objectives): 237 / 286docs/evals/cmmc-evidence-map.md, test split, 26 Sep 2026
- Partial or missing objectives called present (false present): 0 / 121docs/evals/cmmc-evidence-map.md, test split
- Requirements called fully evidenced that are not: 0 / 59docs/evals/cmmc-evidence-map.md, test split
- Present objectives called present (recall): 119 / 165docs/evals/cmmc-evidence-map.md, test split
- Quotes verbatim from the input: 446 / 446docs/evals/cmmc-evidence-map.md, test split
- Objective status, self-hosted direct route (same 286 test objectives): 235 / 286, false present 0 / 121docs/evals/cmmc-evidence-map.md, self-hosted test run