Skip to content
decosa

56 · Compliance and trust · Software and AI ops · live

Incident notification pack

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • Planted problems caught, held-out scenarios20 / 20test splitn = 20Two scenarios run twice: wrong times, wrong counts, unsupported and contradicted claims, removed elements, stale figures, a late 8-K.
  • Planted problems caught, dev scenarios28 / 28dev (tuned on)n = 28
  • Deadlines right, hand-labelled cases, first run35 / 44test splitn = 44Labelled from the rules by a separate agent; the 9 misses were one bug, fixed, after which 44 / 44 (not independent).
  • Clean sentences held, held-out scenarios1 / 42test splitn = 423 of 42 flagged as partly supported.
  • False contradictions on clean held-out packs2 / 4test splitn = 4One pattern: a later time read as the detection time.
  • Model-drafted sentences held (real errors caught)2 / 108syntheticn = 108Both were times copied from the wrong entry; 9 flagged; 28 / 28 required elements given.

Dataset

Five synthetic incidents written by the building agent (ransomware at a SaaS vendor, a public bucket at a DORA payment institution, an exploited router vulnerability under the CRA; held out: a DDoS on a DNS provider and email compromise at a billing company), each run clean and with planted problems, twice; plus 44 deadline cases hand-labelled by a separate agent.

Caveats

  • Everything is synthetic, written by the same agent that wrote the checker and the prompts; small n (5 scenarios).
  • Plants are single clear errors; legal adequacy and subtle understatement are not measured.
  • The deadline set was used to find and fix a bug, so 44 / 44 after the fix is not held out.
  • The judge varies run to run: the same clean sentence was traced in one run and flagged or held in another.
  • Hand labels are by an AI agent from the rules text, not by counsel.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
23 s
Receipts
16
Model calls
n/a
Tokens
n/a
Cost per run
$0.006

Self-host verification

Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers

A fresh clone of a decosa-api pre-release build (not yet merged to main) into a clean directory, the api image built from docker/api/Dockerfile, compose api service with a named data volume, DECOSA_INCIDENT_SYNTHETIC_ONLY=0, direct route to the already-running local Qwen3.8-27B vLLM (network_mode host instead of starting a second model server). The ransomware sample without the synthetic flag: both planted sentences held, impact missing, 3 contradictions, the NIS2 early warning late, report and timeline record verified, a changed status failed, receipts attested, 3.8 s; the CRA sample with a drafted final report 4.4 s. Torn down after. Model-server startup itself not re-verified.

Rehearsal bundle: incident-notification-pack.zip (5 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Synthetic only on the hosted demo, and every number here comes from our own synthetic incidents.
  • Coverage says a traced sentence addresses an element, not that it says enough.
  • The contradiction check can read a later time as the detection time (seen in 2 of 4 clean held-out packs).
  • Member State bank holidays, national NIS2 formats and other US states are not modelled.
  • The clocks are only as right as the tags: the team marks when it became aware, classified or determined materiality.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Timeline, hash chain, deadline clocks, citation, time and number checks, element coverage, cross-notice consistency, signed pack and sign-off (no model; CPU)decosa-api incident pack (decosa_api/verticals/incident) with the numeric block (decosa_api/verticals/numeric), the dates module (decosa_api/verticals/claims/dates.py) and the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
  • Notice drafting, the grounding judge, and the review call (element coverage and quoted facts)Qwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · clocks and timeline, no GPU (1)
  • Deadlines right, 44 hand-labelled cases (SEC, NIS2, DORA, CRA, California): 44 / 44 after one bug fix; 35 / 44 first rundocs/evals/incident-notification-pack.md, part 1, 26 Sep 2026
Standard · one GPU for the model (hosted demo) (5)
  • Planted problems caught in supplied notices (2 repeats): 28 / 28 dev; 20 / 20 held-out testdocs/evals/incident-notification-pack.md, parts 2-6
  • Unsupported or contradicted claims held: 10 / 10docs/evals/incident-notification-pack.md, parts 2-6
  • Clean sentences held / flagged: 1 / 102 held; 14 / 102 flaggeddocs/evals/incident-notification-pack.md, part 7
  • False contradictions on clean packs: 0 of 6 dev packs; 2 of 4 test packs (one pattern)docs/evals/incident-notification-pack.md, part 7
  • Model-drafted sentences traced / flagged / held (5 scenarios): 97 / 9 / 2 (both held were real errors)docs/evals/incident-notification-pack.md, part 8

How we measure · All tools