Skip to content
decosa

23 · Software and AI ops · live

Open-model migration check

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Not-worse agreement with human experts, both orders (run 1)79.3% (74.4-83.5)test splitn = 300Runs 2 and 3: 77.3% and 78.3%. GPT-4 (MT-Bench's own judge) on the same items: 78.0%. On par, not better.
  • Cohen's kappa vs experts (run 1)0.588test splitn = 300Runs 2 and 3: 0.546, 0.568; GPT-4: 0.562
  • Precision / recall on "worse" (run 1)0.755 / 0.839test splitn = 300Errs on the strict side more than the lenient one.
  • Verdict identical in all three runs, per item84.3%test splitn = 300Temperature 0 on a batched server is not bit-exact; variation lands on close calls.
  • Report-level verdict equal to the experts' labels14 of 15 pairingstest splitn = 15GPT-4: 13 of 15. 14 of 15 pairings kept the same verdict across three runs.
  • Dev agreement, both orders85.0%dev (tuned on)n = 120GPT-4 87.5% on the same dev items

Dataset

lmsys/mt_bench_human_judgments (CC-BY-4.0): expert pairwise votes on first-turn MT-Bench answers from six 2023-era models, with GPT-4's own verdicts as the baseline judge. Split by question: 120 dev items, 300 test items; the prompt was written once, run once on dev, not changed, and test was run three times.

Caveats

  • Measures only the free-text judge; the structured scorers are deterministic code covered by unit tests.
  • MT-Bench answers are 2023-era and general-purpose; agreement on a team's own task should be checked against a few of their own labels.
  • Human tie votes are noisy (about a quarter of items); not-worse folds them into "not worse".
  • First-turn answers only; no multi-turn conversations.
  • Self-preference when the judge model judges its own answers is not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
3.6 s
Receipts
30
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers

Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. 20 synthetic tickets: schema valid 20 of 20, a verdict with reasons; the record verifies and recomputes, and a flipped score fails at that entry with the numbers that no longer follow named. Key minting with the admin secret works.

Rehearsal bundle: migration-check.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • The hosted numbers are for the 20-ticket JSON sample. The 20-question MT-Bench free-text sample (70 calls with the judge) took 19 s on a quiet GPU and 3-7 minutes while the shared GPU was busy (25 Sep 2026).
  • Agreement with your current model is not correctness; send human labels where you have them.
  • Latency in the report is measured on a shared GPU through the gateway; your own deployment will differ.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)decosa-api migration module (decosa_api/verticals/migration)AGPL-3.0-or-later
  • Candidate and free-text judgeQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · structured prompts on a 24-32 GB card (self-host) (2)
  • Structured scoring (exact, label, JSON schema, fields): deterministic code, covered by unit tests; no model to measuretests/test_migration.py
  • Gemma-4-31B as a candidate: not measured yetnot measured yet
Standard · Qwen3.8-27B as candidate and judge (hosted demo) (4)
  • MT-Bench test, agreement with human experts on 'is the candidate worse?' (both orders): 79.3% (95% CI 74.4-83.5), κ 0.588; runs 2 and 3: 77.3%, 78.3%docs/evals/migration-check.md, 300 held-out pairs
  • Same items, GPT-4 as judge (published MT-Bench verdicts): 78.0% (73.0-82.3), κ 0.562docs/evals/migration-check.md
  • Agreement without ties (judge vs experts): 89.5% (GPT-4: 88.0%)docs/evals/migration-check.md
  • Report verdict unchanged across 3 runs / matches the experts' verdict: 14 of 15 pairings / 14 of 15 (GPT-4: 13 of 15)docs/evals/migration-check.md
Best · DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host) (1)
  • As a migration candidate: not measured yetnot measured yet
Wanted · the largest open candidates (1)
  • agreement with expert labels, same protocol: not measured yet

How we measure · All tools