Typed-judgment API (24): eval
Run on 25 Sep 2026 on our server (one RTX PRO 6000, shared with other evals the whole time) with scripts/judgment_eval.py.
Model: Qwen3.8-27B NVFP4 (nvidia/Qwen3.8-27B-NVFP4 at revision 482ca0f3…, weights root 2fafb365…1091894), vLLM 0.29.0,
thinking off. Prompt: core.SYSTEM, sha256 d5a22aed85926cbe…. Raw outputs and metrics.json are in ~/.cache/judgment-eval/
on our server (not committed: they contain dataset text).
Data (all public, permissive or attribution licences)
| Set | What it tests | Dev (fit only) | Test (held out, run once) | Licence |
|---|---|---|---|---|
BoolQ (google/boolq validation) |
yes/no with a passage | 300 | 1,000 other questions | CC BY-SA 3.0 |
MMLU (cais/mmlu) |
one of four, no passage | 300 from validation | 1,000 from test | MIT |
SummEval (mteb/summeval) |
eval mode, rubric: 1-5 score per criterion vs expert mean | 10 articles (640 judgments) | 25 articles (1,600) | MIT |
MT-Bench human judgments (lmsys/mt_bench_human_judgments) |
eval mode, pairwise vs expert votes | 150 votes | 600 votes | CC BY 4.0 |
Splits are fixed by seed 24 in build. Nothing was tuned on a test split. The prompt was written once before the dev run
and not changed afterwards.
Methods
- logprobs: one temperature-0 call; the answer token's top-20 log-probabilities renormalised over the allowed answers, then temperature-scaled. T fit on dev by NLL: 1.7 (yes/no), 1.7 (choice); for scores, T = 3.0 chosen by the best mean Spearman on the SummEval dev articles. Needs the model server's logprobs: direct route (self-host) only today.
- samples: the temperature-0 answer plus k answers sampled at temperature 1 (answer line only), vote counts smoothed with a Dirichlet prior: alpha 0.3 (yes/no), 0.2 (choice), 0.2 (score, by dev Spearman). The eval took n=8 samples per item from vLLM in one request and used the first k; the service makes k seeded calls instead (same distribution).
- single: the model's stated confidence (0-100) mapped through an isotonic table fit on dev.
Calibration and accuracy (held-out test)
Confidence = the probability the API returns for its own answer. ECE: 15 equal-width bins. Brier: squared error of that probability against right/wrong. AUROC: how well the probability separates right from wrong answers.
| Set (n) | Method | Accuracy | ECE | Brier | AUROC | ECE before calibration |
|---|---|---|---|---|---|---|
| BoolQ (1,000) | logprobs | 90.7% | 0.023 | 0.069 | 0.857 | 0.059 |
| samples k=2 | 0.016 | 0.078 | 0.625 | |||
| samples k=4 (hosted default) | 0.023 | 0.079 | 0.651 | 0.029 | ||
| samples k=8 | 0.022 | 0.077 | 0.702 | |||
| single (stated) | 0.076 | 0.088 | 0.703 | 0.077 | ||
| MMLU (1,000) | logprobs | 83.4% | 0.026 | 0.102 | 0.863 | 0.058 |
| samples k=2 | 0.057 | 0.115 | 0.739 | |||
| samples k=4 | 0.053 | 0.114 | 0.768 | 0.123 | ||
| samples k=8 | 0.048 | 0.111 | 0.799 | |||
| single (stated) | 0.124 | 0.152 | 0.602 | 0.132 |
Every test answer was in the allowed set (format 100%), and the answer token's logprobs were readable for 100%.
Through the hosted route (service stage: 200 BoolQ and 200 MMLU test questions sent to POST /judgment/batch on a
pre-release server behind the model gateway, method samples, k=4, seed 0, reasons off):
| Set | Accuracy | ECE | Brier | AUROC | Same answer as the direct run | Cost per 1,000 |
|---|---|---|---|---|---|---|
| BoolQ (200) | 90.5% | 0.032 | 0.077 | 0.656 | 97.0% | $0.61 |
| MMLU (200) | 84.0% | 0.028 | 0.121 | 0.683 | 96.5% | $0.58 |
Reading: logprobs are the only method that ranks right against wrong answers well (AUROC ~0.86). Samples give a well-calibrated average but coarse probabilities; the stated confidence is almost always 95-100 and carries little information. The gateway passthrough (below) would give the hosted API the logprobs numbers at a fifth of the calls.
On dev I also tried a two-feature logistic map (vote share plus stated confidence): Brier on BoolQ dev fell from 0.072 to 0.067 at k=4, less on MMLU. Not shipped: the gain is small and it adds a model to maintain.
Eval mode
Rubric (SummEval, 25 held-out articles, 16 summaries each). The expected score (probability-weighted level) against the expert mean, Spearman correlation computed per article and averaged, as in G-Eval:
| Method | Coherence | Consistency | Fluency | Relevance | Mean Spearman | Mean Kendall |
|---|---|---|---|---|---|---|
| logprobs | 0.665 | 0.480 | 0.452 | 0.505 | 0.525 | 0.420 |
| samples k=4 | 0.610 | 0.478 | 0.402 | 0.473 | 0.491 | 0.407 |
| single | 0.636 | 0.573 | 0.421 | 0.450 | 0.520 | 0.466 |
For orientation only: the G-Eval paper reports a mean Spearman of 0.514 for its GPT-4 judge on all 100 articles, with a different prompt and scoring, so the numbers are not directly comparable.
Pairwise (MT-Bench, 600 held-out expert votes, turns 1 and 2). Each pair judged in both orders, distributions averaged:
| Agreement with ties (n=600) | Without ties | Both orders same winner | |
|---|---|---|---|
| Ours, logprobs | 64.8% | 81.3% (459) | 92% |
| Ours, samples k=4 | 64.5% | 81.0% (457) | 92% |
| GPT-4 pair judge, same 600 rows (from the dataset) | 65.2% | 83.6% (384) | n/a |
The probability of the pairwise winner is not well calibrated (ECE 0.17 with logprobs): treat pairwise probabilities as a ranking signal, and do not threshold them without checking on your own data.
Determinism (temperature-0 repeat agreement)
Pinned weights and a pinned prompt rule out the model being replaced. They do not make a busy inference server bit-reproducible: requests are batched with other traffic, and a different batch changes floating-point results.
| Check | BoolQ | MMLU |
|---|---|---|
| Direct route, the temperature-0 call repeated (1,000 each; first run 16 in flight, repeat 4 in flight) | 99.0% same answer | 95.7% |
| identical generated text (answer, confidence and reason) | 45.7% | 21.8% |
| median / p99 largest change in any answer's logprob probability | 0.001 / 0.19 | 0.005 / 0.33 |
| answers with calibrated p ≥ 0.9 that changed | 0 of 757 | 0 of 517 |
| answers with p ≥ 0.75 that changed | 0 of 909 | 1 of 667 |
| Hosted route, whole pipeline (answer + 4 seeded samples) run twice, 200 each | 98.5% same answer | 97.5% |
| identical distributions | 93% | 83% |
On the hosted route none of the 182 answers at p ≥ 0.9 changed; every change was at p ≤ 0.77 in one of the runs. So the API
marks answers under 0.75 as near_tie. In the recorded demo runs, the two changes (an urgency score split 53/37 between two
levels, and "does the lead have 5 support staff" when the page lists 4 plus 3 open roles) were both flagged near-ties.
What would make re-runs exact: a batch-invariant serving mode (vLLM has one on recent builds; whether it supports this NVFP4 checkpoint and what it costs in throughput has not been checked), or a dedicated, unbatched judge instance.
Cost per 1,000 judgments (gateway list price for qwen3.8-27b: $0.30 per million prompt tokens, $1.50 per million generated)
Measured on the hosted route, where the gateway counted about 380 prompt tokens per call for a ~300-token context:
| Setting | Calls per question | Cost per 1,000 |
|---|---|---|
| single, reasons off | 1 | ~$0.12 |
| single, with a reason (~37 generated tokens) | 1 | ~$0.17 |
| samples k=4 (hosted default), reasons off | 5 | $0.58-0.61 (measured) |
| logprobs, self-hosted | 1 | GPU time only |
Longer contexts scale the prompt part linearly. usage.cost_usd in every response is computed from the tokens actually used.
Gateway logprobs passthrough (proposed, not applied)
The gateway's ChatRequest has no logprobs fields and _sku_message_payload forwards a fixed list (top_p, stop,
seed), so logprobs never reach vLLM; the upstream response is otherwise returned as is. docs/proposals/gateway-logprobs.patch adds logprobs and top_logprobs (≤ 20) to the request, forwards them on
non-streaming calls only, strips logprobs from any response that did not ask for them, and adds a logprobs_sha256 receipt field
(unsigned for now; putting it in the signed receipt needs a receipt-format change). Reverse-tunnel providers do not
carry logprobs in this patch. With the patch applied to a copy, the gateway's receipt, SKU, route-profile and
allowance-routing tests pass (40), plus docs/proposals/test_gateway_logprobs_passthrough.py. Once it is live, set
DECOSA_JUDGMENT_GATEWAY_LOGPROBS=1 and method: auto switches to logprobs on the hosted route.
Not measured
- Labels (multi-label) are one yes/no call per label and use the yes/no calibration; no multi-label dataset was run.
- Score calibration beyond ranking (SummEval has no single right level).
- Domains other than these benchmarks. Calibration varies by domain; users should check on their own labelled data.
- The Gemma-4-26B-A4B judge planned on page 26: not served.
Reproduce
python scripts/judgment_eval.py build --data ~/data/typed-judgment --work ~/.cache/judgment-eval
python scripts/judgment_eval.py run --work ~/.cache/judgment-eval --split dev
python scripts/judgment_eval.py run --work ~/.cache/judgment-eval --split test
python scripts/judgment_eval.py repeat --work ~/.cache/judgment-eval --split test --sets boolq,mmlu --concurrency 4
DECOSA_EVAL_KEY=dk_... python scripts/judgment_eval.py service --work ~/.cache/judgment-eval --tag a # then --tag b
python scripts/judgment_eval.py score --work ~/.cache/judgment-eval --calibration decosa_api/verticals/judgment/data/calibration.json
The data directory holds the parquet files from Hugging Face converted to jsonl (boolq-validation, mmlu-validation,
mmlu-test, summeval, mtbench-human, mtbench-gpt4).