Eval: incident notification pack (56)
Run 26 Sep 2026 on our server against on a pre-release build, hosted gateway route, Qwen3.8-27B
(NVFP4), temperature 0. Every incident is synthetic and was written by the agent that built the checker. Result files
are in docs/evals/incident-notification-pack/. A drafting and checking aid: nothing here measures whether a notice is
legally sufficient, and nothing here says "compliant".
What was measured
- Deadline accuracy (code, no model): 44 cases across SEC 8-K, NIS2, DORA, CRA and California.
- Grounding: planted sentences the log does not support or contradicts, in supplied notices.
- Times and numbers: planted wrong clock times (including a local time written as UTC or CEST) and wrong counts and amounts.
- Required-element coverage: a required element removed from a notice.
- Contradictions between notices: a stale figure in one notice while another uses the revised one.
- Missed deadlines: a submission logged after its deadline.
- False alarms on clean packs: held and flagged sentences, elements reported missing, contradictions and late clocks where nothing was planted.
- Model drafts: every notice of each scenario drafted by the model, then checked.
Sets
- Clock cases (
clock-cases.json): 44 cases labelled by hand from the rules text by a separate agent that did not see the clock code; each case carries its reasoning. It covers US federal holidays (Labor Day, Columbus Day, Thanksgiving, Christmas, New Year, Juneteenth, 4 July observed on 3 July), late-night Eastern determinations, weekend determinations, both DST changes, month ends (31 Jan, 31 May), DORA's 24-hour cap and late classification, DORA weekend relief allowed and not allowed, NIS2 and CRA weekend extensions, and California's 30-day and 15-day periods. 17 cases include a submission and a labelled status. - dev (3 scenarios: ransomware at a SaaS vendor with SEC 8-K + NIS2; a public bucket at a DORA payment institution with a California notice; an actively exploited router vulnerability under the CRA). The review prompt, the clean notice wording and the 8-K "impact" word rule were adjusted while looking at these. Two scenario-data bugs were fixed here (a clean log whose moved submission time re-numbered the entries, and a clean DORA log whose intermediate report was late).
- test (2 scenarios written after that and run as written: a DDoS on a DNS provider under NIS2; email compromise and wire fraud at a billing company with SEC 8-K + California). Nothing was changed after seeing these results.
- Each scenario was run as a planted pack and as a clean pack, twice (repeats), through
POST /incident/packwith supplied notices only (drafted notices are scored separately).scripts/eval_incident.py.
Results
1. Deadlines (44 hand-labelled cases)
| First run | After the fix | |
|---|---|---|
| All fields right (due, extended due, status, late by) | 35 / 44 | 44 / 44 |
| Due time right | 35 / 44 | 44 / 44 |
| Status right (17 cases with a submission) | 11 / 17 | 17 / 17 |
All 9 first-run misses were one bug: a later report (NIS2 final, DORA intermediate/final, CRA incident final, the
California AG copy) said "not started" when the earlier report's submission was logged but its own start was not
tagged. It was fixed and a regression test added; the second run is therefore not independent of this set. The SEC
cases (14, holidays and late-night Eastern included) were right on the first run. clock-results-first-run.json,
clock-results.json.
2-6. Planted problems in supplied notices (2 repeats each)
| Kind | dev | test |
|---|---|---|
| Wrong clock time held as a time mismatch | 2 / 2 | 4 / 4 |
| Wrong count or amount held as a number mismatch | 4 / 4 | 4 / 4 |
| Claim the log contradicts, held | 6 / 6 | 2 / 2 |
| Claim no cited entry makes, held | - | 2 / 2 |
| Required element removed, reported not given | 8 / 8 | 4 / 4 |
| Stale figure, listed as a contradiction between notices | 2 / 2 | 2 / 2 |
| Late submission, clock shows late | 6 / 6 | 2 / 2 |
| All planted problems | 28 / 28 | 20 / 20 |
Unsupported and contradicted claims together: 10 of 10 held (6 dev, 4 test); every one was held by the grounding judge with the matching verdict. Times and numbers: 14 of 14, all in code. In the test set, the sentence that carried the stale figure was itself held once (the judge read "first detected at 13:58 UTC" against a latency alert as unsupported, which is defensible); no other unplanted sentence was held in a planted pack.
7. False alarms on clean packs
| dev (6 packs, 60 sentences) | test (4 packs, 42 sentences) | |
|---|---|---|
| Sentences held | 0 | 1 (of 42) |
| Sentences flagged (partly supported) | 11 | 3 |
| Required ("must") elements reported not given | 0 | 0 |
| Contradictions between notices | 0 | 2 (2 packs) |
| Late clocks | 0 | 0 |
- The held sentence (test): "On November 4, 2026, [the company] confirmed that an unauthorized party had accessed the mailbox of its controller since November 2, 2026", which the log supports; the judge called it contradicted in one repeat and partly supported in the other.
- The two contradictions are one pattern in both repeats of the DNS scenario: the review call read "At 14:20 UTC, 38% of queries were failing" as the detection time and compared it with the early warning's "detected at 14:05 UTC". A false alarm; not fixed, because it came from the test set (a guard that would have fixed it could only be checked on that same set, so it stays documented; revisit with new held-out scenarios).
- Flags: 9 of the 11 dev flags are the CRA notice (4 or 5 of its 6 sentences each run). The judge treats small wording differences ("diag endpoint" for "/cgi-bin/diag", "being actively exploited" for "treated as actively exploited") as partial. Flags do not remove a sentence; they ask for a look.
8. Model drafts (5 scenarios, every notice drafted, one run)
- 108 drafted sentences: 97 traced, 9 flagged, 2 held. Both held sentences were real errors by the model, caught in code: it gave the NIS2 early warning the time of another entry (09:30 UTC instead of 19:05 UTC) and gave the ransom demand the early warning's time and date. Without the check both would have gone into a draft 8-K.
- Required ("must") elements given by a traced sentence: 28 of 28.
- 1 contradiction between drafted notices (email-compromise scenario: "confirmed on 4 November" read as detection against "from 2 November"), a false alarm of the same kind as above.
- Not labelled beyond this: the 9 flags were read by the building agent and are small unsupported details (a word such as "fraudulent", "as of the final report").
Speed and cost (hosted route, shared gateway)
- Supplied-notice packs: median 11.8 s (dev) and 24.1 s (test) per pack; the streamed samples took 21-32 s while other agents shared the gateway. Self-hosted on the direct route: 3.8 s (ransomware sample) and 4.4 s (CRA sample with a drafted final report).
- Ransomware sample: 16 model calls, 13,291 tokens, about $0.006; the DORA sample with a drafted notice: 22 calls, about $0.008 (tokens and cost from the gateway receipts).
- During the eval the gateway was down for about 25 minutes (20:20-20:45 UTC, an unrelated incident). Those runs were discarded and everything was re-run afterwards; the reported runs had no failed model calls.
Expected properties of the sample run (rehearsal bundle rehearsal/incident-notification-pack/)
- The 8-K sentence "at 04:12 UTC" is held as
time_mismatch, and the detail names 03:12 UTC (entry E2). - The NIS2 sentence "No personal data was exfiltrated" is held.
- The 8-K's element
impactismissing. contradictionsincludessystems_affected(212 vs 240 hosts).- The NIS2 early-warning clock is
lateby7 h 0 min; the 8-K is due2026-09-22T21:30:00Z(17:30 EDT on the fourth business day). - The signed report, the log hash and the hash-chained timeline verify; a report with its status changed does not.
Limits (read these before quoting the numbers)
- Small and synthetic. 5 scenarios, 48 planted problems over 4 repeats, 102 clean sentences; one author wrote the incidents, the plants and the checker. This shows the mechanisms work on these cases, not accuracy on real incidents.
- Plants are single clear errors. Subtle misstatements, omissions of degree ("the notice understates the impact") and legal adequacy are not measured.
- The clock set was used to find and fix a bug, so 44/44 is not a held-out number; 35/44 is the first-run result.
- Hand labels by an AI agent, not by counsel. Legal ambiguities (California has no weekend roll-forward in the text; whether Regulation 1182/71 extends EU day and month periods) are shown as conventions, not settled.
- Run-to-run variation: the same clean sentence can be traced in one run and flagged or held in another (seen twice).
- Coverage is a mapping, not adequacy: "given" means a traced sentence addresses the element, not that it says enough.
- Code changes after the eval runs, not re-measured by it: plain-text log parsing (tags after extra pipes,
#comment lines, unknown-tag warnings, grouped duplicate warnings), a size limit on JSON logs, a pack with anynot_checkedsentence can no longer end asneeds_review, and a clearer message when a local time is written as UTC. The eval used JSON logs and had no failed calls, so none of these touch its numbers.
Verdict
The checks do what they say on these cases: every planted wrong time, count, contradicted or unsupported claim, removed element, stale figure and late submission was caught (48 of 48), and the model's own drafting errors were caught in code. The costs are a strict judge (14 of 102 clean sentences flagged, 1 held) and a contradiction check that can mistake a later time for the detection time. Ready for a pilot with breach counsel on their own (redacted) past incidents; not evidence of accuracy on real logs until that is done.