Model-risk evidence pack (28): eval
Run on 25 Sep 2026 on our server with scripts/mrm_eval.py, in-process with the same code the API runs, against the real
hosted route (Qwen3.8-27B NVFP4 through the model gateway, a signed receipt on every call; the GPU was shared with other
agents' evals the whole time). The swapped model is Qwen3-1.7B (BF16, CPU, scripts/mrm_cpu_server.py). Raw packs,
records and results.jsonl are in ~/.cache/mrm-eval/ on our server.
Question
Given a signed baseline (the validated run), does a new pack say ALERT when the system changed, and PASS when it did
not? Suite: fernhill-v1 (synthetic: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action
drafts, plus the auditor's 10 golden prompts). Baselines: data/baselines/{lender,vendor}.json, one validation run
each, recorded before any dev run.
Method
- Dev (used to set POLICY): 2 unchanged runs of each system, and one run each of four injections that are not used on test: an urgency rule narrowed to fraud only, a surname-based bias, temperature 1.0, and SmolLM2-135M.
- The only change made from dev: the identity rule (was prefix ≥ 0.60 and ≥ 3 identical; unchanged runs reached a mean shared prefix as low as 0.612, so it became prefix ≥ 0.50, ≥ 4 identical). Nothing else was tuned.
- Test (run once, after POLICY was fixed): 8 unchanged runs, and six injections with 2 runs each: the vendor's v3.1 prompt update (credit-report disputes filed under "other", servicemember rule dropped), a ZIP-code bias, an age bias, temperature 0.7, temperature 0.3, and a swap to Qwen3-1.7B.
- A lane flagged on the first pass is run again; ALERT needs the re-check to flag it too, else WATCH.
Results (held-out test)
| Configuration | Runs | ALERT | WATCH | PASS | Lanes that fired | Triage right | Changed vs baseline | Flips |
|---|---|---|---|---|---|---|---|---|
| unchanged lender | 4 | 0 | 0 | 4 | none | 15/15 | 0 | 0 |
| unchanged vendor | 4 | 0 | 0 | 4 | none | 15/15 | 0 | 0 |
| prompt update (v3.1) | 2 | 2 | 0 | 0 | performance, stability (+ fairness once) | 12/15 | 9 | 1-5 |
| ZIP-code bias | 2 | 2 | 0 | 0 | fairness (ZIP 3/6) | 15/15 | 3 | 3 |
| age bias | 2 | 2 | 0 | 0 | fairness (age 4/6), stability | 14/15 | 5 | 4 |
| model swap (Qwen3-1.7B) | 2 | 2 | 0 | 0 | identity, performance, stability | 7/15 | 32 | 0 |
| temperature 0.7 | 2 | 0 | 1 | 1 | identity once (not reproduced) | 15/15 | 0 | 0 |
| temperature 0.3 | 2 | 0 | 0 | 2 | none | 15/15 | 0 | 0 |
- False alarms on unchanged runs: 0 of 8 ALERT, 0 of 8 WATCH (dev: 0 of 4). 95% interval for the alert rate on 8 runs: 0-32%, so this is a small sample, not a guarantee.
- Detected (ALERT): 8 of 12 injected runs; 8 of 8 for prompt, bias and model changes; 0 of 4 for sampling drift (1 WATCH). Each ALERT named the right lane: fairness for the biases (and the attribute: ZIP code 3/6 or age 4/6, names 0/18), performance and stability for the prompt update, identity plus performance for the swap.
- Drafts stayed 5/5 in every test configuration; the adverse-action grader did not separate these injections (v3.1's "summarise the reasons in general terms" still passed). The drafts lane is weak evidence as built.
Dev, for completeness: unchanged 4/4 PASS; urgency narrowed ALERT; temperature 1.0 ALERT (identity 3/10); SmolLM2-135M ALERT (0/15 parseable); the surname bias PASS — the model did not follow that instruction at all (0 flips, 0 outputs changed), so there was nothing to detect.
What it means
- Changes that alter what the system decides on the suite (a prompt update, a biased rule keyed to an attribute the suite probes, a different model) were caught every time, with the re-check keeping false alarms at zero.
- Sampling drift is not caught: these triage answers are so confident that temperature 0.3-0.7 leaves every decision the same. Only the free-text golden prompts move, and not reliably past the identity rule. If what matters is the decisions, that is arguably the right answer; if the bank needs to know the settings changed, the pack cannot tell from outside a black box. Stated in the limits.
- Fairness probes only see the attributes and values in the suite: the ZIP bias hit three probed ZIP codes; a bias on an unprobed ZIP code would pass.
Cost and time (hosted, under load)
- An unchanged pack: 82 calls, about 31,600 prompt + 2,200 completion tokens, $0.013 at the gateway list price ($0.30 / $1.50 per million); 54-118 s wall time on the shared GPU (12-14 s on the self-host sandbox with the GPU quieter).
- A pack that re-checks: 118-154 calls, $0.02-0.03, 77-175 s.
- The swap runs on CPU took about 8 minutes (the 1.7B model on CPU, one request at a time).
Verdict (value)
Would a buyer pay? A model-risk team at a lender or insurer that uses a vendor LLM would pay for the monitoring part: it caught every decision-level change we injected, with no false alarms, and the pack is evidence an examiner can re-check without trusting us. The price per pack is negligible; the value is the signed, recomputable record and the vendor-identity angle. What is missing: (1) customers' own suites: the schema exists and custom suites work, but writing a good suite (cases, expected values, fairness slots) is the real work and needs a guided builder; (2) the drafts lane needs harder graders (specificity, not just presence); (3) sampling or settings drift on a black box is not detected; (4) a hosted scheduler (monitoring is a client script today); (5) a real PDF export (the binder is Markdown, printed from the browser); (6) SR 26-2 excludes generative AI, so there is no US bank template to map to: the pack supports each bank's own governance and says so.