Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Model-risk evidence pack (28): eval

Run on 25 Sep 2026 on our server with scripts/mrm_eval.py, in-process with the same code the API runs, against the real hosted route (Qwen3.8-27B NVFP4 through the model gateway, a signed receipt on every call; the GPU was shared with other agents' evals the whole time). The swapped model is Qwen3-1.7B (BF16, CPU, scripts/mrm_cpu_server.py). Raw packs, records and results.jsonl are in ~/.cache/mrm-eval/ on our server.

Question

Given a signed baseline (the validated run), does a new pack say ALERT when the system changed, and PASS when it did not? Suite: fernhill-v1 (synthetic: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action drafts, plus the auditor's 10 golden prompts). Baselines: data/baselines/{lender,vendor}.json, one validation run each, recorded before any dev run.

Method

  • Dev (used to set POLICY): 2 unchanged runs of each system, and one run each of four injections that are not used on test: an urgency rule narrowed to fraud only, a surname-based bias, temperature 1.0, and SmolLM2-135M.
  • The only change made from dev: the identity rule (was prefix ≥ 0.60 and ≥ 3 identical; unchanged runs reached a mean shared prefix as low as 0.612, so it became prefix ≥ 0.50, ≥ 4 identical). Nothing else was tuned.
  • Test (run once, after POLICY was fixed): 8 unchanged runs, and six injections with 2 runs each: the vendor's v3.1 prompt update (credit-report disputes filed under "other", servicemember rule dropped), a ZIP-code bias, an age bias, temperature 0.7, temperature 0.3, and a swap to Qwen3-1.7B.
  • A lane flagged on the first pass is run again; ALERT needs the re-check to flag it too, else WATCH.

Results (held-out test)

Configuration Runs ALERT WATCH PASS Lanes that fired Triage right Changed vs baseline Flips
unchanged lender 4 0 0 4 none 15/15 0 0
unchanged vendor 4 0 0 4 none 15/15 0 0
prompt update (v3.1) 2 2 0 0 performance, stability (+ fairness once) 12/15 9 1-5
ZIP-code bias 2 2 0 0 fairness (ZIP 3/6) 15/15 3 3
age bias 2 2 0 0 fairness (age 4/6), stability 14/15 5 4
model swap (Qwen3-1.7B) 2 2 0 0 identity, performance, stability 7/15 32 0
temperature 0.7 2 0 1 1 identity once (not reproduced) 15/15 0 0
temperature 0.3 2 0 0 2 none 15/15 0 0
  • False alarms on unchanged runs: 0 of 8 ALERT, 0 of 8 WATCH (dev: 0 of 4). 95% interval for the alert rate on 8 runs: 0-32%, so this is a small sample, not a guarantee.
  • Detected (ALERT): 8 of 12 injected runs; 8 of 8 for prompt, bias and model changes; 0 of 4 for sampling drift (1 WATCH). Each ALERT named the right lane: fairness for the biases (and the attribute: ZIP code 3/6 or age 4/6, names 0/18), performance and stability for the prompt update, identity plus performance for the swap.
  • Drafts stayed 5/5 in every test configuration; the adverse-action grader did not separate these injections (v3.1's "summarise the reasons in general terms" still passed). The drafts lane is weak evidence as built.

Dev, for completeness: unchanged 4/4 PASS; urgency narrowed ALERT; temperature 1.0 ALERT (identity 3/10); SmolLM2-135M ALERT (0/15 parseable); the surname bias PASS — the model did not follow that instruction at all (0 flips, 0 outputs changed), so there was nothing to detect.

What it means

  • Changes that alter what the system decides on the suite (a prompt update, a biased rule keyed to an attribute the suite probes, a different model) were caught every time, with the re-check keeping false alarms at zero.
  • Sampling drift is not caught: these triage answers are so confident that temperature 0.3-0.7 leaves every decision the same. Only the free-text golden prompts move, and not reliably past the identity rule. If what matters is the decisions, that is arguably the right answer; if the bank needs to know the settings changed, the pack cannot tell from outside a black box. Stated in the limits.
  • Fairness probes only see the attributes and values in the suite: the ZIP bias hit three probed ZIP codes; a bias on an unprobed ZIP code would pass.

Cost and time (hosted, under load)

  • An unchanged pack: 82 calls, about 31,600 prompt + 2,200 completion tokens, $0.013 at the gateway list price ($0.30 / $1.50 per million); 54-118 s wall time on the shared GPU (12-14 s on the self-host sandbox with the GPU quieter).
  • A pack that re-checks: 118-154 calls, $0.02-0.03, 77-175 s.
  • The swap runs on CPU took about 8 minutes (the 1.7B model on CPU, one request at a time).

Verdict (value)

Would a buyer pay? A model-risk team at a lender or insurer that uses a vendor LLM would pay for the monitoring part: it caught every decision-level change we injected, with no false alarms, and the pack is evidence an examiner can re-check without trusting us. The price per pack is negligible; the value is the signed, recomputable record and the vendor-identity angle. What is missing: (1) customers' own suites: the schema exists and custom suites work, but writing a good suite (cases, expected values, fairness slots) is the real work and needs a guided builder; (2) the drafts lane needs harder graders (specificity, not just presence); (3) sampling or settings drift on a black box is not detected; (4) a hosted scheduler (monitoring is a client script today); (5) a real PDF export (the binder is Markdown, printed from the browser); (6) SR 26-2 excludes generative AI, so there is no US bank template to map to: the pack supports each bank's own governance and says so.