Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Typed-judgment API (24): eval

Run on 25 Sep 2026 on our server (one RTX PRO 6000, shared with other evals the whole time) with scripts/judgment_eval.py. Model: Qwen3.8-27B NVFP4 (nvidia/Qwen3.8-27B-NVFP4 at revision 482ca0f3…, weights root 2fafb365…1091894), vLLM 0.29.0, thinking off. Prompt: core.SYSTEM, sha256 d5a22aed85926cbe…. Raw outputs and metrics.json are in ~/.cache/judgment-eval/ on our server (not committed: they contain dataset text).

Data (all public, permissive or attribution licences)

Set What it tests Dev (fit only) Test (held out, run once) Licence
BoolQ (google/boolq validation) yes/no with a passage 300 1,000 other questions CC BY-SA 3.0
MMLU (cais/mmlu) one of four, no passage 300 from validation 1,000 from test MIT
SummEval (mteb/summeval) eval mode, rubric: 1-5 score per criterion vs expert mean 10 articles (640 judgments) 25 articles (1,600) MIT
MT-Bench human judgments (lmsys/mt_bench_human_judgments) eval mode, pairwise vs expert votes 150 votes 600 votes CC BY 4.0

Splits are fixed by seed 24 in build. Nothing was tuned on a test split. The prompt was written once before the dev run and not changed afterwards.

Methods

  • logprobs: one temperature-0 call; the answer token's top-20 log-probabilities renormalised over the allowed answers, then temperature-scaled. T fit on dev by NLL: 1.7 (yes/no), 1.7 (choice); for scores, T = 3.0 chosen by the best mean Spearman on the SummEval dev articles. Needs the model server's logprobs: direct route (self-host) only today.
  • samples: the temperature-0 answer plus k answers sampled at temperature 1 (answer line only), vote counts smoothed with a Dirichlet prior: alpha 0.3 (yes/no), 0.2 (choice), 0.2 (score, by dev Spearman). The eval took n=8 samples per item from vLLM in one request and used the first k; the service makes k seeded calls instead (same distribution).
  • single: the model's stated confidence (0-100) mapped through an isotonic table fit on dev.

Calibration and accuracy (held-out test)

Confidence = the probability the API returns for its own answer. ECE: 15 equal-width bins. Brier: squared error of that probability against right/wrong. AUROC: how well the probability separates right from wrong answers.

Set (n) Method Accuracy ECE Brier AUROC ECE before calibration
BoolQ (1,000) logprobs 90.7% 0.023 0.069 0.857 0.059
samples k=2 0.016 0.078 0.625
samples k=4 (hosted default) 0.023 0.079 0.651 0.029
samples k=8 0.022 0.077 0.702
single (stated) 0.076 0.088 0.703 0.077
MMLU (1,000) logprobs 83.4% 0.026 0.102 0.863 0.058
samples k=2 0.057 0.115 0.739
samples k=4 0.053 0.114 0.768 0.123
samples k=8 0.048 0.111 0.799
single (stated) 0.124 0.152 0.602 0.132

Every test answer was in the allowed set (format 100%), and the answer token's logprobs were readable for 100%.

Through the hosted route (service stage: 200 BoolQ and 200 MMLU test questions sent to POST /judgment/batch on a pre-release server behind the model gateway, method samples, k=4, seed 0, reasons off):

Set Accuracy ECE Brier AUROC Same answer as the direct run Cost per 1,000
BoolQ (200) 90.5% 0.032 0.077 0.656 97.0% $0.61
MMLU (200) 84.0% 0.028 0.121 0.683 96.5% $0.58

Reading: logprobs are the only method that ranks right against wrong answers well (AUROC ~0.86). Samples give a well-calibrated average but coarse probabilities; the stated confidence is almost always 95-100 and carries little information. The gateway passthrough (below) would give the hosted API the logprobs numbers at a fifth of the calls.

On dev I also tried a two-feature logistic map (vote share plus stated confidence): Brier on BoolQ dev fell from 0.072 to 0.067 at k=4, less on MMLU. Not shipped: the gain is small and it adds a model to maintain.

Eval mode

Rubric (SummEval, 25 held-out articles, 16 summaries each). The expected score (probability-weighted level) against the expert mean, Spearman correlation computed per article and averaged, as in G-Eval:

Method Coherence Consistency Fluency Relevance Mean Spearman Mean Kendall
logprobs 0.665 0.480 0.452 0.505 0.525 0.420
samples k=4 0.610 0.478 0.402 0.473 0.491 0.407
single 0.636 0.573 0.421 0.450 0.520 0.466

For orientation only: the G-Eval paper reports a mean Spearman of 0.514 for its GPT-4 judge on all 100 articles, with a different prompt and scoring, so the numbers are not directly comparable.

Pairwise (MT-Bench, 600 held-out expert votes, turns 1 and 2). Each pair judged in both orders, distributions averaged:

Agreement with ties (n=600) Without ties Both orders same winner
Ours, logprobs 64.8% 81.3% (459) 92%
Ours, samples k=4 64.5% 81.0% (457) 92%
GPT-4 pair judge, same 600 rows (from the dataset) 65.2% 83.6% (384) n/a

The probability of the pairwise winner is not well calibrated (ECE 0.17 with logprobs): treat pairwise probabilities as a ranking signal, and do not threshold them without checking on your own data.

Determinism (temperature-0 repeat agreement)

Pinned weights and a pinned prompt rule out the model being replaced. They do not make a busy inference server bit-reproducible: requests are batched with other traffic, and a different batch changes floating-point results.

Check BoolQ MMLU
Direct route, the temperature-0 call repeated (1,000 each; first run 16 in flight, repeat 4 in flight) 99.0% same answer 95.7%
identical generated text (answer, confidence and reason) 45.7% 21.8%
median / p99 largest change in any answer's logprob probability 0.001 / 0.19 0.005 / 0.33
answers with calibrated p ≥ 0.9 that changed 0 of 757 0 of 517
answers with p ≥ 0.75 that changed 0 of 909 1 of 667
Hosted route, whole pipeline (answer + 4 seeded samples) run twice, 200 each 98.5% same answer 97.5%
identical distributions 93% 83%

On the hosted route none of the 182 answers at p ≥ 0.9 changed; every change was at p ≤ 0.77 in one of the runs. So the API marks answers under 0.75 as near_tie. In the recorded demo runs, the two changes (an urgency score split 53/37 between two levels, and "does the lead have 5 support staff" when the page lists 4 plus 3 open roles) were both flagged near-ties.

What would make re-runs exact: a batch-invariant serving mode (vLLM has one on recent builds; whether it supports this NVFP4 checkpoint and what it costs in throughput has not been checked), or a dedicated, unbatched judge instance.

Cost per 1,000 judgments (gateway list price for qwen3.8-27b: $0.30 per million prompt tokens, $1.50 per million generated)

Measured on the hosted route, where the gateway counted about 380 prompt tokens per call for a ~300-token context:

Setting Calls per question Cost per 1,000
single, reasons off 1 ~$0.12
single, with a reason (~37 generated tokens) 1 ~$0.17
samples k=4 (hosted default), reasons off 5 $0.58-0.61 (measured)
logprobs, self-hosted 1 GPU time only

Longer contexts scale the prompt part linearly. usage.cost_usd in every response is computed from the tokens actually used.

Gateway logprobs passthrough (proposed, not applied)

The gateway's ChatRequest has no logprobs fields and _sku_message_payload forwards a fixed list (top_p, stop, seed), so logprobs never reach vLLM; the upstream response is otherwise returned as is. docs/proposals/gateway-logprobs.patch adds logprobs and top_logprobs (≤ 20) to the request, forwards them on non-streaming calls only, strips logprobs from any response that did not ask for them, and adds a logprobs_sha256 receipt field (unsigned for now; putting it in the signed receipt needs a receipt-format change). Reverse-tunnel providers do not carry logprobs in this patch. With the patch applied to a copy, the gateway's receipt, SKU, route-profile and allowance-routing tests pass (40), plus docs/proposals/test_gateway_logprobs_passthrough.py. Once it is live, set DECOSA_JUDGMENT_GATEWAY_LOGPROBS=1 and method: auto switches to logprobs on the hosted route.

Not measured

  • Labels (multi-label) are one yes/no call per label and use the yes/no calibration; no multi-label dataset was run.
  • Score calibration beyond ranking (SummEval has no single right level).
  • Domains other than these benchmarks. Calibration varies by domain; users should check on their own labelled data.
  • The Gemma-4-26B-A4B judge planned on page 26: not served.

Reproduce

python scripts/judgment_eval.py build  --data ~/data/typed-judgment --work ~/.cache/judgment-eval
python scripts/judgment_eval.py run    --work ~/.cache/judgment-eval --split dev
python scripts/judgment_eval.py run    --work ~/.cache/judgment-eval --split test
python scripts/judgment_eval.py repeat --work ~/.cache/judgment-eval --split test --sets boolq,mmlu --concurrency 4
DECOSA_EVAL_KEY=dk_... python scripts/judgment_eval.py service --work ~/.cache/judgment-eval --tag a   # then --tag b
python scripts/judgment_eval.py score  --work ~/.cache/judgment-eval --calibration decosa_api/verticals/judgment/data/calibration.json

The data directory holds the parquet files from Hugging Face converted to jsonl (boolq-validation, mmlu-validation, mmlu-test, summeval, mtbench-human, mtbench-gpt4).