Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: CMMC / NIST 800-171 evidence map (65)

Run 26 Sep 2026 on our server, Qwen3.8-27B through the shared model gateway (other workloads were using it), branch the pre-release branch. Script: scripts/eval_cmmc.py. Results: docs/evals/cmmc-evidence-map/{dev,test}.json (every row) and {dev,test}-outputs.json (every objective's quotes, typed judgment and grounding label).

What is measured

Synthetic small defence contractors from decosa_api/verticals/cmmc/synth.py: 8 NIST SP 800-171 Rev 2 requirements each, drawn from a pool of 14 (51 NIST SP 800-171A objectives), with an SSP statement and the artefacts behind it, and at most one planted gap per requirement: evidence missing (the SSP says implemented, nothing sent), planned, stale (a point-in-time export 400-800 days old), out of scope (from a system not in the SSP's scope list), draft policy, one objective shown only in part, one objective contradicted (a setting off), one objective's line dropped. Every objective has a label (present, partial, missing) computed from which lines survive, with the rule the product uses: the SSP counts only for objectives that ask for something to be defined, identified or specified; an artefact with an issue makes an objective partial at most; a contradicted objective is missing.

  • Dev: 10 companies, 80 requirements, 299 objectives, phrasing A (a Microsoft GCC High stack). Used while writing the prompts and rules.
  • Test: 10 companies, 80 requirements, 286 objectives, phrasing B (on-premises AD, Duo, Wazuh, Nessus, Sophos, FortiGate): every evidence line is worded differently from dev. Generated with fixed seeds before any test run, and run once, after the dev work was frozen, with the production settings (no extra signals).

Metrics per objective: 3-class accuracy; false "present" (labelled partial or missing, called present: the risky direction), as a share of all objectives labelled partial or missing; precision and recall of "present"; citation validity (every quote is a line of the input, verbatim; for objectives called present or partial, a quote is one of the planted evidence lines for that objective); and the requirement-level roll-up.

Results

Dev (final rules, run 3) Test (held out, run once)
Objectives, 3-class accuracy 238 / 299 (79.6%) 237 / 286 (82.9%)
False "present" (labelled partial or missing) 3 / 102 (2.9%) 0 / 121
Precision of "present" 152 / 155 119 / 119
Recall of "present" 152 / 197 119 / 165 (72%)
Not concluded (a call failed) 0 0
Requirements, roll-up accuracy 62 / 80 64 / 80
Requirements called fully evidenced that are not 0 / 51 0 / 59
Quotes verbatim from the input 438 / 438 446 / 446
Present or partial calls quoting a planted evidence line 223 / 223 204 / 209
Model calls (receipts) per company 97.5 75.6
Seconds per company (8 requirements, 3 in parallel, shared gateway) 176.5 236.7

Test by planted gap (objectives right / total): planned 36/36, scope 28/29, stale 53/56, draft 4/4, no_evidence 11/12, contra 25/33, drop 25/34, partial 7/10, none 48/72.

Test confusion (label -> call): present -> present 119, partial 30, missing 16; partial -> partial 60, missing 1; missing -> missing 58, partial 2.

Where it errs, it mostly errs down. 47 of the 49 test errors call an objective less evidenced than its label; the other 2 call a missing objective partial (never present). The biggest groups: 3.1.1[b] (service accounts listed, "processes acting on behalf of users identified") read as partial 6 times; 3.1.1[e] and 3.1.12[a] missed as no evidence 5 and 5 times (the FortiGate and "Log on as a service" lines were not cited); MFA for privileged local and network access (3.5.3[b], [c]) read as partial. We did not change anything after seeing the test split.

Self-hosted route (direct, log-probabilities)

Self-hosted, the typed judgment reads the model's log-probabilities (one call) instead of 3 votes. Checked in a fresh clone built into a container (a clean directory, DECOSA_CMMC_SYNTHETIC_ONLY=0) against the local Qwen3.8-27B vLLM, then torn down.

Dev Test (run once, after the change below)
Objectives, 3-class accuracy 140 / 299 as first run; 232 / 299 on replay with the change 235 / 286 (82.2%)
False "present" 0 / 102 0 / 121
Recall of "present" 54 / 197 first run; 146 / 197 on replay 117 / 165
Requirements called fully evidenced that are not 0 / 51 0 / 59
Quotes verbatim 459 / 459 450 / 450
Model calls per company / seconds per company 39.2 / 24.1 s

The one change, made on dev: a typed "present" on this route was demoted whenever the judgment engine flagged a near tie (p < 0.75), which kept only 54 of 197 present objectives. The rule is now p >= 0.5 on every route, as on the gateway. (dev-selfhost.json is the first run; the replay is scripts/eval_cmmc.py's decide() over its saved outputs.) The gateway numbers above did not change: the vote-based method never sets that flag.

The rehearsal bundle passed 21 of 21 checks in 19.8 s on this box, and the smoke module passed in 5.3 s with 18 attested receipts; a set not marked synthetic and one with a CUI banner were accepted, as they should be in real-data mode.

How the rules were set (dev only)

Three dev runs; statuses are replayable from the saved signals (eval_cmmc.py dev --replay, and mapper.decide).

  1. First prompts: 215 / 299, false present 0 / 102, but 60 present objectives called partial.
  2. Changes: the typed judgment's context grouped by artefact with the SSP and the system scope as context; "a plan to do it later" moved from partial to missing in both prompts; the grounding judge's "partial" accepted (it is literal about wording: "designated locations"); four template lines rewritten where the model was right and the label was not (MFA only on servers, one of two external connections verified, an SSP line that named service accounts). 239 / 299, false present 7 / 102.
  3. Fixes for those 7: objectives like "vulnerabilities are identified" (a when-clause) or "flaws are identified within the specified time" were wrongly treated as documentation objectives, so the SSP alone counted; the rule now excludes when / within / and-implemented. A dropped objective no longer leaves a hollow artefact (title and filler lines). "Present" now needs the majority of the typed votes (p >= 0.5); an unsure present is partial. 238 / 299, false present 3 / 102. The 3 remaining dev false presents quote a line from another requirement's artefact that arguably does bear on the objective (Intune blocking personal devices for 3.1.1[f]; the workload-identity list for 3.1.1[b]; a KEV rescan for 3.14.1[b]): the labels only count a requirement's own artefacts.

Expected properties of the demo sample (harbor-precision)

Checkable on any run (rehearsal/cmmc-evidence-map/expected.json has them all):

  1. 3.13.11 is missing (the SSP says planned); none of 3.1.8, 3.3.1, 3.5.3, 3.5.7, 3.10.3, 3.14.2 is fully evidenced.
  2. Artefact A4 is flagged stale, A5 draft, A9 out of scope; the test-run certificate T1 verifies; the screenshot S1 is transcribed by the vision model.
  3. The POA&M row for 3.10.3 says it can never be on a POA&M; 3.13.11's row is conditional.
  4. No estimate or score is shown (9 of 110 requirements).
  5. The signed record verifies at /record/verify and fails once a requirement's status is changed.

Limits (read these before quoting a number)

  • Everything is synthetic, from templates written by the same agent that wrote the prompts and rules: this shows the mechanisms work on text shaped like evidence, not accuracy on real SSPs. No RPO or assessor labelled anything.
  • 14 of the 110 requirements are in the pool. Families like CM, MA, PS and RA are not exercised.
  • Labels count only a requirement's own artefacts; real evidence is shared across requirements, and the model sometimes (rightly) uses it.
  • Recall of "present" is 72% on test: by design every doubt goes down, so a preparer will re-check some objectives that are in fact evidenced.
  • Screenshots: only the one bundled sample was read (the vision step works; its accuracy on real console screenshots is not measured). Scanned PDFs are not read at all.
  • No SPRS number is validated: the estimate needs all 110 requirements in one run, which the hosted demo does not allow.