Metrics
Trust
Accuracy
How well each tool does. For every tool: the numbers from its own eval, with the dataset, what was held out and the caveats; the latest nightly end-to-end check on the hosted API with latency, receipts and cost; and whether a fresh self-host setup has been run. Small, honest numbers, not a leaderboard.
Machine-readable: /api/metrics.json. How the numbers are made: how we measure.
- What this page proves
- Each tool's numbers on its own test set, with the data it used (real or public, or synthetic, labelled), what was held out, and the date.
- What it doesn't
- How a tool will do on your documents. A synthetic test set is not real-world performance; run your own sample first.
- Tools
- 86
- every workspace on the site
- With eval numbers
- 85
- a measured quality eval of their own
- Held out or test split
- 69
- scored on cases kept apart from tuning
- Self-host verified
- 73
- fresh setup run end to end
- Rehearsal bundles
- 85
- mock-data zip you can run yourself
Every tool
Loading the nightly status from the API…
Showing 86 of 86 tools
- Agent flight recorderSoftware and AI ops · Compliance and trustEval headline
Held-out agent runs reaching the expected outcome: 8/8
held-out / test splitheld out+6 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 5.3 s · $0.003 / run8 receipts - Audio drama and narrated story studioFilm, TV and games · Creative and mediaEval headline
Radio scripts: speaker right (code alone): 100% (389/389)
held-out / test splittest split+6 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 28 s · $0.001 / run1 receipt - Auto F&I disclosure recordSales and marketing · Finance and insuranceEval headline
Per-check accuracy, test B (held out), run 1: 31/32
held-out / test splitheld out+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 3.3 s · $0.003 / run10 receipts - Capture audit evidence from your admin screensCompliance and trust · Software and AI opspreviewEval headline
Quarterly verdicts right on held-out consoles (no model): 33 / 33
not held outtest split+8 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 7.5 s · $0 / run0 receipts - Certificate request checkFinance and insurancepreviewEval headline
Status right on found requirements: 109 of 120 (0.91)
not held outtest split+3 more29 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 33 s · $0.015 / run31 receipts - Check their briefLegalEval headline
Non-existent citations caught (strict): 24 / 31 (77%)
held-out / test splitheld out+8 more28 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 8.3 s · $0.013 / run7 receipts - Eval headline
Deficient responses caught on unseen real cases (RECAP test, dockets never seen in dev): 174 / 235 (74%)
held-out / test splittest split+7 more28 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 4.8 s · $0.006 / run10 receipts - Citation and claim checker for papersScience and researchEval headline
Planted wrong-paper citations flagged: 12 of 12
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 9.6 s · $0.003 / run4 receipts - Claim denial appeal packetHealthcare · Finance and insuranceEval headline
Recommendation right, fresh held-out set: 36 of 40
held-out / test splittest split+8 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 13 s · $0.014 / run21 receipts - Clinical AI assurance monitorHealthcare · Compliance and trustEval headline
Invented fact caught as an error: 49 (98%)
held-out / test splittest split+9 more28 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 28 s · $0.021 / run30 receipts - CMMC / NIST 800-171 evidence mapCompliance and trust · Public sectorEval headline
Objective status, 3 classes: 237 / 286
held-out / test splittest split+8 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 150 s · $0.008 / run33 receipts - Collections and servicing call QAFinance and insurance · Compliance and trustEval headline
Per-rule accuracy, test run 1: 81/83
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 4.2 s · $0.003 / run9 receipts - Consented creator dubbingFilm, TV and games · Creative and mediaEval headline
Dub tracks exactly the video's length (sample count): 11 of 11
not held outsynthetic+6 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 74 s · $0.003 / run43 receipts - CSR number-to-table verifierScience and research · HealthcareEval headline
Planted number errors caught: 92 / 96
held-out / test splittest split+5 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 6.5 s · $0.004 / run6 receipts - Decosa StudioCreative and media · Personal and familyEval headline
Quality not measured yet
not held out23 Sep 2026Nightly checkPartial 25 SepQALatency and costp50 n/a · cost n/a1 receipt - Eval headline
Contradiction finder: planted conflicts found (held-out set): 9 / 9 in both runs
held-out / test splittest split+4 more24 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 40 s · $0.014 / run48 receipts - Device complaint MDR triageHealthcare · Compliance and trustEval headline
Reportable complaints called not reportable: 6 of 96 (6.3%)
held-out / test splittest split+7 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 12 s · $0.002 / run5 receipts - Disclosed UGC adsSales and marketing · Creative and mediapreviewEval headline
Watermark survival on the two example ads (metadata stripped / CRF 28 re-encode / 50% resize): receipt id recovered in all 6 checks (6/7 to 7/7 frames vote)
not held outsynthetic+1 more24 Sep 2026Nightly checkPartial 25 SepQALatency and costp50 6.7 s · $0.003 / run6 receipts - Disclosed virtual stagingCreative and media · Sales and marketingEval headline
Planted edits caught (check): 38 / 40
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 15 s · $0.002 / run3 receipts - Editorial-control ledgerCreative and media · Compliance and trustEval headline
Wording-only paraphrases with no claim change flagged (held out): 24 / 24
held-out / test splitheld out+5 more25 Sep 2026Nightly checkPartial 25 SepQALatency and costp50 59 s · $0.002 / run2 receipts - Endpoint auditorSoftware and AI ops · Compliance and trustEval headline
Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b: fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences
not held outsynthetic+6 more24 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 40 s · $0.002 / run22 receipts - EU trial lay summary with number groundingHealthcare · Science and researchEval headline
Planted errors caught, drafts with citations (code checks): 555 / 577 (96.2%)
held-out / test splittest split+9 more27 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 8.2 s · $0.001 / run3 receipts - Expert-to-SOPField and trades · Any industryEval headline
Planted on-screen steps found: 21 / 25
held-out / test splittest split+6 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 20 s · $0.016 / run16 receipts - Family interview filmCreative and media · Film, TV and gamesEval headline
Shown quotes inside her own turn, at the right time: 46 / 46
not held outdev (tuned on)+9 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 41 s · $0.009 / run11 receipts - Field reportsField and tradesEval headline
Expected issues in the report (grounded report + safety sweep + claim check): 40/41 (98%); before 38/41 (93%)
not held outsynthetic+3 more23 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 26 s · $0.026 / run28 receipts - Filing pre-flightLegalEval headline
Fake citations caught as problem (strict): 89%
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 12 s · $0.015 / run15 receipts - Filing tie-out and MD&A groundingFinance and insurance · Compliance and trustEval headline
False flags on untouched held-out 10-K MD&As (frozen rules): 11 in 6,039 figures (1.8 per 1,000)
held-out / test splittest split+8 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 9.5 s · <$0.001 / run1 receipt - Fill a form from your papersAny industry · Finance and insurancepreviewEval headline
PDF forms: wrong answers / answers written: 1 / 491 (0.20%)
held-out / test splittest split+9 more28 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 14 s · $0.003 / run3 receipts - GPSR listing packCompliance and trust · Sales and marketingEval headline
Missing Article 19 elements found, frozen: 15 / 19
held-out / test splittest split+8 more27 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 12 s · $0.004 / run37 receipts - Green-claims substantiation checkCompliance and trustEval headline
Verdict accuracy, v2 on test2 (held out, first run): 43/44 (98%)
held-out / test splitheld out+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 8.6 s · $0.013 / run26 receipts - Grounding checkAny industry · Compliance and trustEval headline
Unsupported-sentence precision / recall, default gate: 0.593 / 0.556
held-out / test splitheld out+4 more24 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 17 s · $0.002 / run5 receipts - HCC evidence file and RADV defenceHealthcare · Finance and insuranceEval headline
Verdict accuracy, 3 classes: 189 / 200
held-out / test splittest split+9 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 24 s · $0.003 / run15 receipts - Honest product imagerySales and marketing · Compliance and trustpreviewEval headline
Planted misrepresentations flagged: 48 / 48
not held outtest split+6 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 13 s · $0.002 / run10 receipts - Incident notification packCompliance and trust · Software and AI opsEval headline
Planted problems caught, held-out scenarios: 20 / 20
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 23 s · $0.006 / run16 receipts - Injury demand readerFinance and insurance · LegalpreviewEval headline
Respond-by date exact: 7 of 7
held-out / test splittest split+4 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 161 s · $0.020 / run23 receipts - Insurance claims-file conduct packFinance and insurance · Compliance and trustEval headline
Planted problems found, all checks: 52 of 54 (precision 0.98, recall 0.96)
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 10 s · $0.004 / run5 receipts - Interview themesScience and researchpreviewEval headline
Words credited to the wrong speaker (12 public-domain interviews): 0.05% (22 of 41,624)
not held outtest split+6 more29 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 69 s · $0.099 / run62 receipts - Label consistency across PI, SmPC, CCDS and cartonHealthcare · Compliance and trustEval headline
Planted drifts caught: 31 / 31
held-out / test splittest split+6 more27 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 2.3 s · $0.001 / run3 receipts - Live translationAny industryEval headline
FLORES-200 devtest en→es: chrF++ / BLEU / COMET-22: 55.3 / 29.9 / 87.2
held-out / test splitheld out+4 more23 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 2.3 s · $0.003 / run18 receipts - M&A due-diligence red flagsLegal · Finance and insuranceEval headline
Planted red flags caught: 20 / 20
held-out / test splittest split+6 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 93 s · $0.003 / run3 receipts - Made-for-kids content pre-flightCreative and media · Compliance and trustEval headline
Made for kids or not (full check): 39 / 40
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 7.2 s · $0.004 / run15 receipts - Medical chronology with page citesLegal · HealthcareEval headline
Planted events found: 189 / 198
held-out / test splittest split+9 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 25 s · $0.004 / run8 receipts - Medicare sales-call recordHealthcare · Sales and marketingEval headline
Per-rule accuracy, test B run 1 (never used to change anything): 40/40
held-out / test splitheld out+6 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 7.5 s · $0.006 / run15 receipts - Mix cue sheetMusic · Creative and mediaEval headline
Song starts within 10 s, with your track files: 105/105 (92/105 within 5 s, 24/105 within 1 s)
held-out / test splittest split+9 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 7.0 s · $0 / run0 receipts - Model-risk evidence packFinance and insurance · Compliance and trustEval headline
False alarms on unchanged runs (ALERT / WATCH): 0 of 8 / 0 of 8
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 62 s · $0.013 / run82 receipts - Music video from your trackMusic · Film, TV and gamesEval headline
Lyric lines within 0.3 s / 1 s of human timing: 81.8% / 87.7%
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 15 s · $0.002 / run7 receipts - Music video starring youMusic · Creative and mediapreviewEval headline
Planned cuts found on the beat in the exports: 27 / 27
not held outsynthetic+8 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 88 s · $0.029 / run5 receipts - No Surprises Act IDR packet and eligibility screenHealthcare · Finance and insuranceEval headline
Eligibility verdict right from documents, held out: 61 of 61
held-out / test splittest split+6 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 52 s · $0.020 / run23 receipts - Notes from your own jottingsHealthcareEval headline
Risk and safety statements carried, fresh blind split #3 (verbatim rule): 113 / 117
held-out / test splitheld out+9 more29 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 25 s · $0.009 / run24 receipts - Open-model migration checkSoftware and AI opsEval headline
Not-worse agreement with human experts, both orders (run 1): 79.3% (74.4-83.5)
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 3.6 s · $0.003 / run30 receipts - Our story filmCreative and media · Film, TV and gamespreviewEval headline
Planned cuts found on the beat in the exports: 30 / 30
not held outsynthetic+8 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 143 s · $0.041 / run5 receipts - Eval headline
Supported-or-not call agrees with the labels: 93.4%
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 7.8 s · $0.020 / run13 receipts - Payer audit responseHealthcare · Compliance and trustEval headline
Weak claims flagged, new blind BCBSM letter (29 Sep): 9 / 9
held-out / test splitheld out+9 more29 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 60 s · $0.018 / run18 receipts - Pharmacovigilance intakeHealthcare · Compliance and trustEval headline
Missing minimum criteria caught: 17 / 20
held-out / test splitheld out+9 more27 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 7.1 s · $0.003 / run4 receipts - Prior-auth pre-check and packetHealthcarepreviewEval headline
Decision right, latest fresh set: 14 of 16
held-out / test splittest split+5 more28 Sep 2026Nightly checkPartial 28 SepQALatency and costp50 41 s · $0.014 / run - Private code assistantSoftware and AI opsEval headline
Coding benchmark, 5 core tasks, direct API (Qwen3.8-27B, vLLM NVFP4 + MTP): 478 / 500
not held outtest split+3 more23 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 4.3 s · <$0.001 / run1 receipt - Eval headline
Privileged vs not, model call: accuracy / precision / recall: 83.9% / 82.3% / 79.7%
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 121 s · $0.031 / run100 receipts - Eval headline
Memo fact recall (blind grader): 89.3%
held-out / test splittest split+5 more28 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 148 s · $0.017 / run44 receipts - Eval headline
Clause detection vs CUAD labels, v1 mapping (like for like): precision / recall: 0.77 / 0.79
held-out / test splitheld out+6 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 28 s · $0.007 / run20 receipts - Promotional-claims pre-checkHealthcare · Compliance and trustEval headline
Planted problems caught (all categories): 20 / 20
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 3.9 s · $0.022 / run14 receipts - Public-records deskPublic sector · LegalEval headline
Personal-data spans redacted (PII recall), final pipeline: 59 of 60 (98.3%)
held-out / test splittest split+6 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 40 s · $0.012 / run41 receipts - Put the visit into your EHR as draftsHealthcarepreviewEval headline
Typed or picked values wrong (read back from the EHR's database): 0 / 419
not held outsynthetic+6 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 90 s · $0.034 / runSelf-hostnot yet - Reg E dispute investigation fileFinance and insurance · Compliance and trustEval headline
Missing notice elements flagged (recall): 0.956 (43/45)
held-out / test splittest split+7 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 38 s · $0.006 / run8 receipts - Report integrityPublic sector · LegalEval headline
Planted additions found: 16/16
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 97 s · $0.004 / run21 receipts - Review reply with patient privacyHealthcare · Sales and marketingEval headline
Healthcare replies that confirm a patient (blind judge): 0 / 40
held-out / test splittest split+4 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 1.8 s · <$0.001 / run1.55 receipts - Rights-cleared music generationMusic · Creative and mediaEval headline
Prompt guard precision (refuse): 100% (48/48)
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 43 s · <$0.001 / run1 receipt - Sample and lyric clearance pre-checkMusic · LegalEval headline
Audio plants found, medium and high confidence: 62 of 108 (57%)
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 6.7 s · <$0.001 / run1 receipt - Sanctions alert disposition recordFinance and insurance · Compliance and trustEval headline
Same-party pairs proposed as false positive (the risky direction): 0 of 1,094
held-out / test splittest split+6 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 8.4 s · <$0.001 / run2 receipts - SAR narrative deskFinance and insurance · Compliance and trustEval headline
Planted wrong numbers caught in reference narratives, first run (no model): 1,598 of 1,600 (99.9%)
held-out / test splittest split+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 25 s · $0.017 / run30 receipts - Script to animaticFilm, TV and games · Creative and mediaEval headline
Shot-list coverage of lines needing a shot, before any code repair: 100% in 10 of 10 runs
not held outsynthetic+5 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 531 s · $0.007 / run16 receipts - Security questionnaire answererCompliance and trustEval headline
Fill precision: 0.970 (64 of 66)
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 63 s · $0.031 / run62 receipts - Eval headline
Cites on the right page, rendered lines: 376 / 376
held-out / test splittest split+7 more29 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 104 s · $0.010 / run23 receipts - Signed lab notebookScience and research · Compliance and trustEval headline
Genuine synthetic exports that verify: 300 / 300
not held outsynthetic+5 more26 Sep 2026Nightly checkPartial 26 SepQALatency and costp50 24 s · $0.001 / run1 receipt - Signed VEX triageSoftware and AI ops · Compliance and trustEval headline
not_affected precision against Canonical's VEX: 154 / 159
held-out / test splittest split+7 more26 Sep 2026Nightly checkPassed 26 SepQALatency and costp50 14 s · $0.014 / run12 receipts - Eval headline
Planted errors found: 43 of 43
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 6.7 s · $0.003 / run5 receipts - Storefront accessibility passSales and marketing · Compliance and trustEval headline
Planted issues caught: 122 / 127
held-out / test splittest split+6 more27 Sep 2026Nightly checkPassed 28 SepQALatency and costp50 1.8 s · <$0.001 / run2 receipts - Structured oral assessmentEducation · HR and recruitingEval headline
Draft level equals the label (exact): 90.8%
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 18 s · $0.007 / run42 receipts - Synthetic-performer disclosure and S&P pre-flightFilm, TV and games · Sales and marketingEval headline
Planted script issues found (full config): 27 of 29 (93%)
held-out / test splittest split+4 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 3.6 s · $0.001 / run5 receipts - Tamper-evident recordPublic sector · Compliance and trustEval headline
Claim check on the synthetic sessions: sentences supported: 14/17 council, 9/9 interview
not held outsynthetic+1 more23 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 17 s · $0.009 / run54 receipts - Tariff classification memoCompliance and trustEval headline
Memo top-1 subheading (6-digit): 145 / 200
held-out / test splittest split+8 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 32 s · $0.004 / run1 receipt - Typed-judgment APIAny industry · Software and AI opsEval headline
BoolQ accuracy / ECE, logprobs (direct route): 90.7% / 0.023
held-out / test splitheld out+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 4.1 s · $0.004 / run35 receipts - Vendor bank-change checkFinance and insurance · Compliance and trustpreviewEval headline
Fraud flagged (score 3 or more): 9 of 9
held-out / test splittest split+4 more29 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 8.1 s · <$0.001 / run1 receipt - Verified end-to-end test runsSoftware and AI ops · Compliance and trustEval headline
Verdict agrees with the scripted Playwright test, first run of each case: 52/52
held-out / test splittest split+5 more25 Sep 2026Nightly checkPassed 25 SepQALatency and costp50 86 s · $0.002 / run6 receipts - Visit copilotHealthcareEval headline
Medical-term miss rate, live ASR (Voxtral Mini 4B Realtime): 8.4%
held-out / test splitheld out+9 more26 Sep 2026Nightly checkPassed 29 SepQALatency and costp50 60 s · $0.070 / run80 receipts - Walkthrough-to-quoteField and tradesEval headline
Planted items found: 33 / 33
held-out / test splittest split+8 more27 Sep 2026Nightly checkPassed 27 SepQALatency and costp50 54 s · $0.020 / run5 receipts - What studies foundScience and research · HealthcarepreviewEval headline
Verdict matches the systematic review's (benefit, harm or no claim): 231 / 289 (80%)
held-out / test splitheld out+7 more30 Sep 2026Nightly checkPassed 30 SepQALatency and costp50 26 s · $0.023 / run32 receipts
Quality not measured yet: Decosa Studio. Their pages say what is and is not checked.
How we measure
Quality evals
Each tool has an eval of its own: cases with planted problems or known answers, scored by a script or a judge model. Where there is a test split, the prompts were tuned on the dev split only and the test split was run once the code was frozen. We show the test number, the dev number where it matters, and the error or false-positive rate next to the hit rate.
Held out, or not
“Held out” means the scored cases were never used to change prompts or code: a test split kept apart, or a public benchmark we did not tune on. Many of our evals are not held out; the page says so. Those numbers show the pipeline works on the cases we wrote, not how well it generalises.
Synthetic data
Most cases are synthetic: written for the eval, often by the same person who wrote the prompts, because real clinical notes, privileged documents or customer calls cannot be published. Synthetic cases tend to be cleaner and more blatant than real ones. Expect real-world numbers to be lower.
Nightly smoke checks
Every night a script calls each tool's hosted API with its own sample input through a dedicated key and checks the answer and its receipts. Pass or fail, latency, receipts, model calls and cost come from that run. Cost is estimated at the gateway list price. A smoke check proves the pipeline runs end to end; it is not a quality score.
Receipts
Model calls on the hosted route return signed receipts: which model ran, on which inputs (by hash), with what output. Receipts prove what ran, not that the answer is right.
What we don't claim
No eval here is an independent audit, a clinical or legal validation, or a certification. Sample sizes are small. Numbers from a proxy task (for example a clinical-notes benchmark shown for a sales tool) are labelled as proxies and are never the headline. Where nothing is measured yet, we say “not measured yet”.
The eval docs live in the decosa-api repository under docs/evals/ (access required). The nightly feed is GET /verify/status on the API; see also receipts and verification, where your data goes and the regulation watch (the laws and rules each tool cites, checked nightly).