23 · Software and AI ops · live
Open-model migration check
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Not-worse agreement with human experts, both orders (run 1)79.3% (74.4-83.5)test splitn = 300Runs 2 and 3: 77.3% and 78.3%. GPT-4 (MT-Bench's own judge) on the same items: 78.0%. On par, not better.
- Cohen's kappa vs experts (run 1)0.588test splitn = 300Runs 2 and 3: 0.546, 0.568; GPT-4: 0.562
- Precision / recall on "worse" (run 1)0.755 / 0.839test splitn = 300Errs on the strict side more than the lenient one.
- Verdict identical in all three runs, per item84.3%test splitn = 300Temperature 0 on a batched server is not bit-exact; variation lands on close calls.
- Report-level verdict equal to the experts' labels14 of 15 pairingstest splitn = 15GPT-4: 13 of 15. 14 of 15 pairings kept the same verdict across three runs.
- Dev agreement, both orders85.0%dev (tuned on)n = 120GPT-4 87.5% on the same dev items
Dataset
lmsys/mt_bench_human_judgments (CC-BY-4.0): expert pairwise votes on first-turn MT-Bench answers from six 2023-era models, with GPT-4's own verdicts as the baseline judge. Split by question: 120 dev items, 300 test items; the prompt was written once, run once on dev, not changed, and test was run three times.
Caveats
- Measures only the free-text judge; the structured scorers are deterministic code covered by unit tests.
- MT-Bench answers are 2023-era and general-purpose; agreement on a team's own task should be checked against a few of their own labels.
- Human tie votes are noisy (about a quarter of items); not-worse folds them into "not worse".
- First-turn answers only; no multi-turn conversations.
- Self-preference when the judge model judges its own answers is not measured.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 3.6 s
- Receipts
- 30
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. 20 synthetic tickets: schema valid 20 of 20, a verdict with reasons; the record verifies and recomputes, and a flipped score fails at that entry with the numbers that no longer follow named. Key minting with the admin secret works.
Rehearsal bundle: migration-check.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- The hosted numbers are for the 20-ticket JSON sample. The 20-question MT-Bench free-text sample (70 calls with the judge) took 19 s on a quiet GPU and 3-7 minutes while the shared GPU was busy (25 Sep 2026).
- Agreement with your current model is not correctness; send human labels where you have them.
- Latency in the report is measured on a shared GPU through the gateway; your own deployment will differ.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)decosa-api migration module (decosa_api/verticals/migration)AGPL-3.0-or-later
- Candidate and free-text judgeQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · structured prompts on a 24-32 GB card (self-host) (2)
- Structured scoring (exact, label, JSON schema, fields): deterministic code, covered by unit tests; no model to measuretests/test_migration.py
- Gemma-4-31B as a candidate: not measured yetnot measured yet
Standard · Qwen3.8-27B as candidate and judge (hosted demo) (4)
- MT-Bench test, agreement with human experts on 'is the candidate worse?' (both orders): 79.3% (95% CI 74.4-83.5), κ 0.588; runs 2 and 3: 77.3%, 78.3%docs/evals/migration-check.md, 300 held-out pairs
- Same items, GPT-4 as judge (published MT-Bench verdicts): 78.0% (73.0-82.3), κ 0.562docs/evals/migration-check.md
- Agreement without ties (judge vs experts): 89.5% (GPT-4: 88.0%)docs/evals/migration-check.md
- Report verdict unchanged across 3 runs / matches the experts' verdict: 14 of 15 pairings / 14 of 15 (GPT-4: 13 of 15)docs/evals/migration-check.md
Best · DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host) (1)
- As a migration candidate: not measured yetnot measured yet
Wanted · the largest open candidates (1)
- agreement with expert labels, same protocol: not measured yet