Skip to content
decosa

28 · Finance and insurance · Compliance and trust · live

Model-risk evidence pack

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • False alarms on unchanged runs (ALERT / WATCH)0 of 8 / 0 of 8test splitn = 895% interval for the alert rate on 8 runs: 0-32%. Dev: 0 of 4.
  • Injected changes detected (ALERT)8 of 12 runstest splitn = 128 of 8 for prompt, bias and model changes; each ALERT named the right lane.
  • Sampling drift detected (temperature 0.3 / 0.7)0 of 4 runstest splitn = 41 WATCH, not reproduced on re-check.
  • Adverse-action drafts lane separating injections5/5 in every configurationtest splitThe drafts grader did not separate any injection: weak evidence as built.
  • Cost per unchanged pack (list price)$0.013test split82 calls; a pack that re-checks: $0.02-0.03

Dataset

Synthetic suite fernhill-v1: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action drafts, plus 10 golden prompts. Dev (2 unchanged runs per system, 4 injections not reused) set the policy; test (8 unchanged runs, 6 different injections x 2 runs) was run once after the policy was fixed.

Caveats

  • Synthetic suite and injections written by the same team that built the pack.
  • Small sample: 8 unchanged runs gives a 0-32% interval on the false-alarm rate.
  • Sampling or settings drift on a black box is not detected.
  • Fairness probes only see the attributes and values in the suite; a bias on an unprobed value would pass.
  • The drafts lane grader checks presence, not specificity.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
62 s
Receipts
82
Model calls
n/a
Tokens
n/a
Cost per run
$0.013

Self-host verification

Verified on 25 Sep 2026: Fresh clone of the pre-release branch, api image built from docker/api/Dockerfile, compose from the assemble prompt (llm service dropped, api on host network pointed at the running Qwen3.8-27B vLLM, named volume).

Verified on 2026-09-25: image builds, service starts healthy, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Lender pack PASS (15/15, 5/5, 82 attested receipts, 12 s); /mrm/verify ok, and a one-word edit in the record fails at that entry; key minting and two mrm_monitor.py runs (baseline, then trend) worked; logs held no case text. Torn down afterwards.

Rehearsal bundle: model-risk-pack.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Detects only what the suite probes: other products, attributes or ZIP codes are not covered.
  • Sampling or settings changes that do not change decisions are not detected (0 of 4 in the eval).
  • The drafts grader checks that listed reasons appear, not how specific they are.
  • Hosted demo sessions allow about three packs (20,000 generated tokens); use an API key for more.
  • Scheduled monitoring is a client script (cron or systemd), not a hosted scheduler.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)decosa-api model-risk module (decosa_api/verticals/mrm)AGPL-3.0-or-later
  • Fixed grader (typed judgments) and the hosted system under testQwen3.8-27B (NVFP4)Apache-2.0
  • Injected problem in the eval: the 'vendor' silently moved to a small modelQwen3-1.7B (BF16, CPU)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · grader on one 32 GB card (self-host) (1)
  • Detection and false alarms on this hardware: not measured yetnot measured yet
Standard · hosted grader and demo systems (Qwen3.8-27B) (4)
  • False alarms on unchanged systems (held-out): 0 of 8 ALERT, 0 of 8 WATCH (95% CI for the alert rate 0-32%)docs/evals/model-risk-pack.md
  • Injected prompt, bias and model changes caught (held-out): 8 of 8 ALERT: vendor prompt update 2/2, ZIP-code bias 2/2, age bias 2/2, swap to Qwen3-1.7B 2/2; each named the right lanedocs/evals/model-risk-pack.md
  • Sampling drift caught (temperature 0.3 and 0.7): 0 of 4 ALERT (1 WATCH): the triage decisions did not change, so the pack did not alarmdocs/evals/model-risk-pack.md
  • Adverse-action drafts lane: 5/5 in every configuration: it did not separate these injections (weak evidence as built)docs/evals/model-risk-pack.md

How we measure · All tools