56 · Compliance and trust · Software and AI ops · live
Incident notification pack
Eval results
Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Planted problems caught, held-out scenarios20 / 20test splitn = 20Two scenarios run twice: wrong times, wrong counts, unsupported and contradicted claims, removed elements, stale figures, a late 8-K.
- Planted problems caught, dev scenarios28 / 28dev (tuned on)n = 28
- Deadlines right, hand-labelled cases, first run35 / 44test splitn = 44Labelled from the rules by a separate agent; the 9 misses were one bug, fixed, after which 44 / 44 (not independent).
- Clean sentences held, held-out scenarios1 / 42test splitn = 423 of 42 flagged as partly supported.
- False contradictions on clean held-out packs2 / 4test splitn = 4One pattern: a later time read as the detection time.
- Model-drafted sentences held (real errors caught)2 / 108syntheticn = 108Both were times copied from the wrong entry; 9 flagged; 28 / 28 required elements given.
Dataset
Five synthetic incidents written by the building agent (ransomware at a SaaS vendor, a public bucket at a DORA payment institution, an exploited router vulnerability under the CRA; held out: a DDoS on a DNS provider and email compromise at a billing company), each run clean and with planted problems, twice; plus 44 deadline cases hand-labelled by a separate agent.
Caveats
- Everything is synthetic, written by the same agent that wrote the checker and the prompts; small n (5 scenarios).
- Plants are single clear errors; legal adequacy and subtle understatement are not measured.
- The deadline set was used to find and fix a bug, so 44 / 44 after the fix is not held out.
- The judge varies run to run: the same clean sentence was traced in one run and flagged or held in another.
- Hand labels are by an AI agent from the rules text, not by counsel.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 23 s
- Receipts
- 16
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.006
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged to main) into a clean directory, the api image built from docker/api/Dockerfile, compose api service with a named data volume, DECOSA_INCIDENT_SYNTHETIC_ONLY=0, direct route to the already-running local Qwen3.8-27B vLLM (network_mode host instead of starting a second model server). The ransomware sample without the synthetic flag: both planted sentences held, impact missing, 3 contradictions, the NIS2 early warning late, report and timeline record verified, a changed status failed, receipts attested, 3.8 s; the CRA sample with a drafted final report 4.4 s. Torn down after. Model-server startup itself not re-verified.
Rehearsal bundle: incident-notification-pack.zip (5 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Synthetic only on the hosted demo, and every number here comes from our own synthetic incidents.
- Coverage says a traced sentence addresses an element, not that it says enough.
- The contradiction check can read a later time as the detection time (seen in 2 of 4 clean held-out packs).
- Member State bank holidays, national NIS2 formats and other US states are not modelled.
- The clocks are only as right as the tags: the team marks when it became aware, classified or determined materiality.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Timeline, hash chain, deadline clocks, citation, time and number checks, element coverage, cross-notice consistency, signed pack and sign-off (no model; CPU)decosa-api incident pack (decosa_api/verticals/incident) with the numeric block (decosa_api/verticals/numeric), the dates module (decosa_api/verticals/claims/dates.py) and the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
- Notice drafting, the grounding judge, and the review call (element coverage and quoted facts)Qwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · clocks and timeline, no GPU (1)
- Deadlines right, 44 hand-labelled cases (SEC, NIS2, DORA, CRA, California): 44 / 44 after one bug fix; 35 / 44 first rundocs/evals/incident-notification-pack.md, part 1, 26 Sep 2026
Standard · one GPU for the model (hosted demo) (5)
- Planted problems caught in supplied notices (2 repeats): 28 / 28 dev; 20 / 20 held-out testdocs/evals/incident-notification-pack.md, parts 2-6
- Unsupported or contradicted claims held: 10 / 10docs/evals/incident-notification-pack.md, parts 2-6
- Clean sentences held / flagged: 1 / 102 held; 14 / 102 flaggeddocs/evals/incident-notification-pack.md, part 7
- False contradictions on clean packs: 0 of 6 dev packs; 2 of 4 test packs (one pattern)docs/evals/incident-notification-pack.md, part 7
- Model-drafted sentences traced / flagged / held (5 scenarios): 97 / 9 / 2 (both held were real errors)docs/evals/incident-notification-pack.md, part 8