{"schema_version":"1","site":"https://decosa.ai","id":"model-risk-pack","num":"28","name":"Model-risk evidence pack","tool_name":"Build a model-risk evidence pack","short":"Model-risk pack","blurb":"Signed evidence that an LLM a bank, insurer or lender relies on still behaves as validated: identity probes, graded test cases, stability against the validated run, counterfactual fairness pairs and a model card, with a re-check before any alert. Validators can recompute every number from the record.","status":"live","labels":{"industry":["finance","compliance-trust"],"job":["attest","review"],"input":["endpoint"],"deploy":["hosted","selfhost"],"status":"live","output":["record","text"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["finance","compliance-trust"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["endpoint-audit","typed-judgment","signed-record"],"models":"Qwen3.8-27B (hosted system and fixed grader); any OpenAI-compatible endpoint under test","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the grader; the pack runner is CPU; the system under test runs wherever it runs","final_artifact":"A signed evidence pack with a Markdown binder, and a hash-chained record a validator can recompute.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per call; signed pack and hash-chained record that recompute","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":62000,"p95_ms":null,"runs":null,"receipts_per_run":82,"cost_per_run_usd":0.013},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh clone of the pre-release branch, api image built from docker/api/Dockerfile, compose from the assemble prompt (llm service dropped, api on host network pointed at the running Qwen3.8-27B vLLM, named volume).","notes":"Verified on 2026-09-25: image builds, service starts healthy, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Lender pack PASS (15/15, 5/5, 82 attested receipts, 12 s); /mrm/verify ok, and a one-word edit in the record fails at that entry; key minting and two mrm_monitor.py runs (baseline, then trend) worked; logs held no case text. Torn down afterwards."},"known_limits":["Detects only what the suite probes: other products, attributes or ZIP codes are not covered.","Sampling or settings changes that do not change decisions are not detected (0 of 4 in the eval).","The drafts grader checks that listed reasons appear, not how specific they are.","Hosted demo sessions allow about three packs (20,000 generated tokens); use an API key for more.","Scheduled monitoring is a client script (cron or systemd), not a hosted scheduler."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"False alarms on unchanged runs (ALERT / WATCH)","value":"0 of 8 / 0 of 8","unit":null,"n":8,"split":"test","note":"95% interval for the alert rate on 8 runs: 0-32%. Dev: 0 of 4."},{"name":"Injected changes detected (ALERT)","value":"8 of 12","unit":"runs","n":12,"split":"test","note":"8 of 8 for prompt, bias and model changes; each ALERT named the right lane."},{"name":"Sampling drift detected (temperature 0.3 / 0.7)","value":"0 of 4","unit":"runs","n":4,"split":"test","note":"1 WATCH, not reproduced on re-check."},{"name":"Adverse-action drafts lane separating injections","value":"5/5 in every configuration","unit":null,"n":null,"split":"test","note":"The drafts grader did not separate any injection: weak evidence as built."},{"name":"Cost per unchanged pack (list price)","value":"$0.013","unit":null,"n":null,"split":"test","note":"82 calls; a pack that re-checks: $0.02-0.03"}],"dataset":"Synthetic suite fernhill-v1: 15 triage cases, 6 fairness bases x 5 one-attribute variants, 5 adverse-action drafts, plus 10 golden prompts. Dev (2 unchanged runs per system, 4 injections not reused) set the policy; test (8 unchanged runs, 6 different injections x 2 runs) was run once after the policy was fixed.","held_out":true,"caveats":["Synthetic suite and injections written by the same team that built the pack.","Small sample: 8 unchanged runs gives a 0-32% interval on the false-alarm rate.","Sampling or settings drift on a black box is not detected.","Fairness probes only see the attributes and values in the suite; a bias on an unprobed value would pass.","The drafts lane grader checks presence, not specificity."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/model-risk-pack"},"stack":{"summary":"Re-runs a fixed test suite against an LLM system you use: your own deployment of an open model, a vendor's black-box API, or any OpenAI-compatible endpoint. Four lanes: identity (the endpoint auditor's golden prompts against the validated run), graded canaries (cases scored exactly, and drafts graded by typed judgments on a fixed open grader), stability (the same inputs as the validated run) and fairness (counterfactual pairs that change only a name, ZIP code or age). A flagged lane is run again, independently, before the pack says alert. You get a signed pack with a model card and a trend line, a Markdown binder, and a hash-chained record of every output with its receipt, from which /mrm/verify recomputes every number. The demo is a fictional credit union's complaint triage and adverse-action reasons.","tagline":"Signed evidence that an LLM your bank, insurer or lending team relies on still behaves as validated, with a record your validators can recompute.","deployment":"hosted-or-self-host","regulatory_note":"Checked 25 Sep 2026; not legal advice. US banks: SR 26-2 (Fed, OCC, FDIC, 17 Apr 2026) replaced SR 11-7 and OCC 2011-12, and its footnote 3 puts generative and agentic AI models outside its scope, leaving their controls to each bank's own risk management and governance. This pack is evidence for that governance; it is not a validation opinion and does not make anyone compliant. Lenders: ECOA / Regulation B (12 CFR 1002.9) requires specific principal reasons for adverse action; the CFPB withdrew Circulars 2022-03 and 2023-03 on 12 May 2025, but the rule stands. Insurers: the NAIC AI model bulletin is adopted in 25 jurisdictions (NAIC map, 1 Apr 2026), insurers stay responsible for vendor AI systems, and a 12-state AI Systems Evaluation Tool pilot ran March to September 2026; NYDFS Circular Letter 7 (2024) asks for testing for unfair discrimination before and after deployment. EU: credit scoring of natural persons is high-risk under the AI Act (Annex III 5(b)); after Regulation (EU) 2026/1744 those obligations apply from 2 Dec 2027. The fairness probes change one attribute at a time among the values in the suite; they do not measure disparate impact on real applicants. Test cases often hold customer data: keep real cases on a self-hosted box. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3-1.7B).","components":[{"id":"runner","role":"Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)","name":"decosa-api model-risk module (decosa_api/verticals/mrm)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; reuses the endpoint auditor's golden prompts and SSRF guard, typed judgments, and the session hash chain","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"qwen","role":"Fixed grader (typed judgments) and the hosted system under test","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, seed 0, thinking off","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"swap","role":"Injected problem in the eval: the 'vendor' silently moved to a small model","name":"Qwen3-1.7B (BF16, CPU)","hf_repo":"Qwen/Qwen3-1.7B","license":"Apache-2.0","params":"1.7B","quant":"BF16","vram_gb":0,"memory_gb_estimate":null,"engine":"transformers on CPU (scripts/mrm_cpu_server.py), greedy, thinking off","receipt_coverage":"none","in_hosted_demo":false,"tiers":[],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · grader on one 32 GB card (self-host)","summary":"The same pinned grader at 4-bit on an RTX 5090; the runner on the same box. Test your own endpoints on your network.","components":["runner","qwen"],"hardware":"1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Detection and false alarms on this hardware","value":"not measured yet","source":"not measured yet"}],"latency_note":"estimate: not run on this card","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Self-hosted calls get receipts signed by your own box (attested), not by the gateway.","hosting":null},{"id":"standard","label":"Standard · hosted grader and demo systems (Qwen3.8-27B)","summary":"What the hosted API runs: every probe of the hosted system and every grader call is receipted by the gateway.","components":["runner","qwen"],"hardware":"1x RTX PRO 6000 96 GB (measured, shared)","quality_evidence":[{"metric":"False alarms on unchanged systems (held-out)","value":"0 of 8 ALERT, 0 of 8 WATCH (95% CI for the alert rate 0-32%)","source":"docs/evals/model-risk-pack.md"},{"metric":"Injected prompt, bias and model changes caught (held-out)","value":"8 of 8 ALERT: vendor prompt update 2/2, ZIP-code bias 2/2, age bias 2/2, swap to Qwen3-1.7B 2/2; each named the right lane","source":"docs/evals/model-risk-pack.md"},{"metric":"Sampling drift caught (temperature 0.3 and 0.7)","value":"0 of 4 ALERT (1 WATCH): the triage decisions did not change, so the pack did not alarm","source":"docs/evals/model-risk-pack.md"},{"metric":"Adverse-action drafts lane","value":"5/5 in every configuration: it did not separate these injections (weak evidence as built)","source":"docs/evals/model-risk-pack.md"}],"latency_note":"measured on a shared GPU through the gateway: one to two minutes for an unchanged pack; longer with a re-check.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Gateway-signed receipt per call, embedded in the signed record.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /mrm/info, /mrm/targets, /mrm/suites/{id}; POST /mrm/runs (SSE or JSON), /mrm/verify, /mrm/render, /mrm/trend. Stores demo packs only."},{"name":"vLLM (grader and hosted system)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"Endpoint auditor golden prompts (vertical 09)","url":null,"license":"Apache-2.0","purpose":"The identity lane: 10 fixed prompts at temperature 0, compared with the validated run and, for documented open weights, with the auditor's signed reference."},{"name":"Typed judgments (vertical 24)","url":null,"license":"Apache-2.0","purpose":"Grades each adverse-action draft: one yes/no per listed reason, one for an added reason or a protected characteristic."},{"name":"Suite fernhill-v1 (synthetic)","url":null,"license":"Apache-2.0","purpose":"15 complaint-triage cases, 6 counterfactual bases x 5 variants (three names, a ZIP code, an age), 5 adverse-action cases. Written for this demo; the credit union is fictional."},{"name":"scripts/mrm_eval.py, docs/evals/model-risk-pack.md","url":null,"license":"Apache-2.0","purpose":"Injects problems (prompt change, bias, drift, model swap) and runs unchanged systems, on separate dev and test splits."},{"name":"scripts/mrm_monitor.py","url":null,"license":"Apache-2.0","purpose":"Scheduled monitoring from cron or a systemd timer: keeps the baseline and each pack, prints the trend, exits 2 on an alert."}],"hardware":[{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: not run for this vertical."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services."},{"tier":"CPU only","fits":true,"notes":"The runner needs no GPU. With the grader on another box (or the hosted API), a CPU box can run packs against any endpoint."}],"latency":[{"lane":"whole pack, unchanged system (82 calls: 10 identity, 15 triage, 36 fairness, 5 drafts, 16 grader), 2-4 in flight, hosted gateway route","typical_ms":62000,"source":"measured on our server 2026-09-25: 54-118 s over 8 held-out runs while the GPU was shared with other evaluation jobs"},{"lane":"whole pack with a re-check of the flagged lanes (118-154 calls)","typical_ms":135000,"source":"measured on our server 2026-09-25: 77-175 s, shared GPU"},{"lane":"whole pack, self-hosted direct route (same GPU, quieter)","typical_ms":13000,"source":"measured in the self-host sandbox 2026-09-25: 12-14 s for 82 calls"}],"benchmark":null,"notes":["Held-out eval: on 8 unchanged runs the pack never alarmed; on 8 runs with a prompt update, a ZIP-code or age bias, or a swap to a small model it said ALERT every time and named the lane and the attribute.","It did not catch sampling turned up to temperature 0.3-0.7: the triage decisions stayed the same. For a black-box vendor the pack sees behaviour, not settings.","Policy was set on a separate dev split with different injections; the only change from dev was the identity threshold, and the test set was run once afterwards.","The stand-in vendor is the same open weights behind a hidden prompt; the swapped-model demo run is a real recorded run against Qwen3-1.7B on CPU, which the hosted box does not keep running."]},"buyer_facts":[{"label":"Typical run cost","value":"A few cents or less per pack at the gateway list price; more with a re-check. Each run shows its own measured cost."},{"label":"Data retention","value":"Hosted: packs for the demo systems (synthetic) are stored; packs for your endpoint or suite are never stored, they go back to you. Self-host: nothing leaves the box."},{"label":"What leaves the box","value":"Hosted: your suite's case texts go to the system under test and, for drafts, to the hosted grader. Self-host: only calls to the endpoint you test."},{"label":"Inputs","value":"A suite of up to 40 classify cases, 10 fairness bases with up to 8 variants, 10 drafts (200 KB); any OpenAI-compatible endpoint, https and public on the hosted API"},{"label":"Evidence","value":"Signed pack (Ed25519) + hash-chained record with every output and receipt; /mrm/verify recomputes every number; Markdown binder"},{"label":"Not","value":"A validation opinion, legal advice, or a disparate-impact study on real applicants"}],"data_handling":{"page":"/data#model-risk-pack","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Hosted: packs for the demo systems (synthetic) are stored; packs for your endpoint or suite are never stored, they go back to you. Self-host: nothing leaves the box.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"The system you test","route":"both","sends":"your-system","what":"Your suite's case texts go to the OpenAI-compatible endpoint you name (https and public on the hosted API).","default":"always","off":null}]},"console":{"href":"/tools/finance/model-risk-pack","input":"mrm","lanes":[{"id":"identity","title":"Identity","kind":"list"},{"id":"performance","title":"Graded canaries","kind":"list"},{"id":"stability","title":"Stability","kind":"list"},{"id":"fairness","title":"Fairness probes","kind":"list"},{"id":"pack","title":"Evidence pack","kind":"json"}],"samples":[{"n":1,"id":"mrm-vendor","title":"Mrm vendor","deep_link":"/tools/finance/model-risk-pack?sample=1&autorun=0"},{"n":2,"id":"mrm-vendor-updated","title":"Mrm vendor updated","deep_link":"/tools/finance/model-risk-pack?sample=2&autorun=0"},{"n":3,"id":"mrm-vendor-biased","title":"Mrm vendor biased","deep_link":"/tools/finance/model-risk-pack?sample=3&autorun=0"},{"n":4,"id":"mrm-vendor-drift","title":"Mrm vendor drift","deep_link":"/tools/finance/model-risk-pack?sample=4&autorun=0"},{"n":5,"id":"mrm-vendor-swapped","title":"Mrm vendor swapped","deep_link":"/tools/finance/model-risk-pack?sample=5&autorun=0"},{"n":6,"id":"mrm-lender","title":"Mrm lender","deep_link":"/tools/finance/model-risk-pack?sample=6&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/model-risk-pack-hosted.md","selfhost":"/prompts/model-risk-pack-selfhost.md","assemble":"/prompts/model-risk-pack-assemble.md","mac":"/prompts/model-risk-pack-mac.md"},"rehearsal":{"bundle":"/samples/model-risk-pack.zip","bundle_url":"https://decosa.ai/samples/model-risk-pack.zip","folder":"/samples/model-risk-pack/","expected":"/samples/model-risk-pack/expected.json","files":["/samples/model-risk-pack/expected.json","/samples/model-risk-pack/inputs/suite.json"],"bytes":2925,"checks":["the run finishes without errors","all 14 calls were made (10 identity probes and 4 cases)","no call failed","at least 3 of the 4 cases are triaged correctly","the planted SCRA request from a servicemember is marked urgent","the planted discrimination allegation is flagged","the pack verdict is pass or watch","the signed pack, its record and the recomputation all verify","a pack with one observed answer changed no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: cases written for the Fernhill suite on 25 Sep 2026; Fernhill Valley Credit Union is fictional and names and complaints are invented. Part of decosa-api, AGPL-3.0-or-later.","about":"A cut-down validation suite (four synthetic member complaints for the fictional Fernhill Valley Credit Union, with the expected queue and flags) run against the documented demo deployment, plus the ten identity probes. The planted cases are a servicemember asking for SCRA protection (must be urgent) and a mortgage applicant describing discrimination (must be flagged). The signed pack and its record must verify and recompute, and a pack with one observed answer changed must not.","run":{"containers":"docker compose exec api python scripts/rehearse.py model-risk-pack","checkout":"python scripts/rehearse.py model-risk-pack --bundle model-risk-pack.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py model-risk-pack"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=model-risk-pack","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/finance/model-risk-pack","json":"/use-cases/model-risk-pack.json","metrics":"/metrics/model-risk-pack","console":"/tools/finance/model-risk-pack","console_sample":"/tools/finance/model-risk-pack?sample=1&autorun=0","stack":"/tools/finance/model-risk-pack#stack","try_live":"/tools/finance/model-risk-pack","watch":"/tools/finance/model-risk-pack","build":"/tools/finance/model-risk-pack#build","self_host":"/tools/finance/model-risk-pack#self-host","prompts":{"hosted":"/prompts/model-risk-pack-hosted.md","selfhost":"/prompts/model-risk-pack-selfhost.md","assemble":"/prompts/model-risk-pack-assemble.md","mac":"/prompts/model-risk-pack-mac.md"}}}