Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: the migration check's free-text judge (vertical 23)

Run 25 Sep 2026 on our server, through the model gateway (every judge call receipted), Qwen3.8-27B NVFP4 on vLLM 0.29.0, temperature 0, seed 7, judge prompt migration-judge-v1 (sha256 in GET /migration/info). Script: scripts/migration_eval.py; per-item results (ids, labels, both judge orders, receipt ids; no answer text) and metrics in docs/evals/migration-check/.

The structured scorers (exact, label, JSON schema and fields) are deterministic code and are covered by tests/test_migration.py, not by this eval. The only model judgment in a report is the free-text judge, so that is what this measures: does it agree with people, and does its verdict hold still when run again?

Data

lmsys/mt_bench_human_judgments (CC-BY-4.0): expert pairwise votes on first-turn MT-Bench answers from GPT-4, GPT-3.5-turbo, Claude-v1, Vicuna-13B, Alpaca-13B and LLaMA-13B, plus GPT-4's own verdicts on the same pairs (gpt4_pair split), which serve as the baseline judge. One item per (question, model pair) with a human majority: 805 items. The "current" side is the closed model when exactly one side is closed, else the first name alphabetically; the human label is read from the candidate's side (better, same, worse).

Split by question: dev = question ids ending in 1, 2 or 3 (24 questions, 120 items sampled), test = the other 56 questions (300 items sampled, seed 0). The prompt was written once and run once on dev (both-orders agreement 85.0% vs GPT-4's 87.5% there); it was not changed after that, and test was run with it three times. Nothing was tuned on test.

Agreement with the human experts (test, 300 items)

"Not worse" is the decision a report makes: the candidate passes unless judged worse.

Judge Not-worse agreement (95% CI) κ Precision / recall on "worse" 3-way agreement Agreement without ties
Qwen3.8-27B, both orders (the default), run 1 79.3% (74.4-83.5) 0.588 0.755 / 0.839 68.0% 89.5% (172 pairs)
same, run 2 77.3% (72.2-81.6) 0.546 0.745 / 0.797 65.9% 89.2% (167)
same, run 3 78.3% (73.3-82.6) 0.568 0.747 / 0.825 66.7% 88.9% (171)
Qwen3.8-27B, one order only (current answer first), run 1 76.3% (71.2-80.8) 0.532 0.694 / 0.902 68.0% 85.9% (199)
GPT-4 (MT-Bench's own judge, published verdicts) 78.0% (73.0-82.3) 0.562 0.739 / 0.832 68.0% 88.0% (183)

The open judge agrees with the experts about as often as GPT-4 did on the same items; the intervals overlap almost entirely, so "on par", not "better". Running both orders costs a second call and buys about 2-3 points of agreement and fewer false "worse" calls (precision 0.69 to 0.75); that is why it is the default. For reference, the MT-Bench paper reports GPT-4 pairwise agreement with experts of 66% with ties and 85% without; our 3-way and no-tie numbers sit in the same range on this sample.

Stability across repeats

  • Per item: the combined verdict was identical in all three runs for 84.3% of items. Temperature 0 on a batched, speculative-decoding server is not bit-exact under load, and the variation lands on close calls.
  • Position: in each run the two orders gave the same outcome on 79.5-79.8% of items; where they disagree the pair counts as "same" (the MT-Bench convention).
  • Report level: grouping the test items by (current model, candidate) pairing gives 15 pairings with 10 or more items. With a 0.8 bar and the report's verdict rule, 14 of 15 pairings got the same verdict in all three runs; the one that moved (LLaMA-13B to Vicuna-13B, 15 of 16) went inconclusive, go, go at the edge of the interval.
  • Against the humans at report level (run 1): the judge's verdict equals the verdict the expert labels would give in 14 of 15 pairings; GPT-4's in 13 of 15. The miss (Claude-v1 to GPT-3.5-turbo) is the judge saying no-go where the experts' labels say inconclusive: it errs on the strict side more often than the lenient one (recall on "worse" is higher than precision).
  • Every report also states a bootstrap stability (share of 1,000 seeded resamples of its examples that reach the same verdict), which tells a reader how close to the bar that particular result is.

Cost and speed of the eval

1,845 judge calls across the three test runs, about 64,000 generated tokens per run. Mean 13-21 s per call with 8 in flight: the GPU was shared with other evals and services all afternoon, so these are not the judge's speed on a quiet card.

Limits

  • MT-Bench answers are 2023-era and general-purpose; production prompts are narrower. Agreement on a team's own task should be checked against a few of their own labels, which the report takes.
  • Human "tie" votes are noisy (about a quarter of items); the not-worse metric folds them into "not worse".
  • Only first-turn answers; no multi-turn conversations.
  • The candidate in the demo is also the judge model. When Qwen3.8-27B judges its own answers there may be self-preference; this eval does not measure that (the judged answers here are other models'). A self-hosted box can point the judge at a different model to cross-check.