28 · Finance and insurance · Compliance and trust · live
Model-risk evidence pack
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- False alarms on unchanged runs (ALERT / WATCH)0 of 8 / 0 of 8test splitn = 895% interval for the alert rate on 8 runs: 0-32%. Dev: 0 of 4.
- Injected changes detected (ALERT)8 of 12 runstest splitn = 128 of 8 for prompt, bias and model changes; each ALERT named the right lane.
- Sampling drift detected (temperature 0.3 / 0.7)0 of 4 runstest splitn = 41 WATCH, not reproduced on re-check.
- Adverse-action drafts lane separating injections5/5 in every configurationtest splitThe drafts grader did not separate any injection: weak evidence as built.
- Cost per unchanged pack (list price)$0.013test split82 calls; a pack that re-checks: $0.02-0.03
Dataset
Synthetic suite fernhill-v1: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action drafts, plus 10 golden prompts. Dev (2 unchanged runs per system, 4 injections not reused) set the policy; test (8 unchanged runs, 6 different injections x 2 runs) was run once after the policy was fixed.
Caveats
- Synthetic suite and injections written by the same team that built the pack.
- Small sample: 8 unchanged runs gives a 0-32% interval on the false-alarm rate.
- Sampling or settings drift on a black box is not detected.
- Fairness probes only see the attributes and values in the suite; a bias on an unprobed value would pass.
- The drafts lane grader checks presence, not specificity.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 62 s
- Receipts
- 82
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.013
Self-host verification
Verified on 25 Sep 2026: Fresh clone of the pre-release branch, api image built from docker/api/Dockerfile, compose from the assemble prompt (llm service dropped, api on host network pointed at the running Qwen3.8-27B vLLM, named volume).
Verified on 2026-09-25: image builds, service starts healthy, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Lender pack PASS (15/15, 5/5, 82 attested receipts, 12 s); /mrm/verify ok, and a one-word edit in the record fails at that entry; key minting and two mrm_monitor.py runs (baseline, then trend) worked; logs held no case text. Torn down afterwards.
Rehearsal bundle: model-risk-pack.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Detects only what the suite probes: other products, attributes or ZIP codes are not covered.
- Sampling or settings changes that do not change decisions are not detected (0 of 4 in the eval).
- The drafts grader checks that listed reasons appear, not how specific they are.
- Hosted demo sessions allow about three packs (20,000 generated tokens); use an API key for more.
- Scheduled monitoring is a client script (cron or systemd), not a hosted scheduler.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)decosa-api model-risk module (decosa_api/verticals/mrm)AGPL-3.0-or-later
- Fixed grader (typed judgments) and the hosted system under testQwen3.8-27B (NVFP4)Apache-2.0
- Injected problem in the eval: the 'vendor' silently moved to a small modelQwen3-1.7B (BF16, CPU)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · grader on one 32 GB card (self-host) (1)
- Detection and false alarms on this hardware: not measured yetnot measured yet
Standard · hosted grader and demo systems (Qwen3.8-27B) (4)
- False alarms on unchanged systems (held-out): 0 of 8 ALERT, 0 of 8 WATCH (95% CI for the alert rate 0-32%)docs/evals/model-risk-pack.md
- Injected prompt, bias and model changes caught (held-out): 8 of 8 ALERT: vendor prompt update 2/2, ZIP-code bias 2/2, age bias 2/2, swap to Qwen3-1.7B 2/2; each named the right lanedocs/evals/model-risk-pack.md
- Sampling drift caught (temperature 0.3 and 0.7): 0 of 4 ALERT (1 WATCH): the triage decisions did not change, so the pack did not alarmdocs/evals/model-risk-pack.md
- Adverse-action drafts lane: 5/5 in every configuration: it did not separate these injections (weak evidence as built)docs/evals/model-risk-pack.md