Skip to content
decosa

24 · Any industry · Software and AI ops · live

Typed-judgment API

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • BoolQ accuracy / ECE, logprobs (direct route)90.7% / 0.023held outn = 1,000AUROC 0.857. Samples k=4 (hosted default): ECE 0.023, AUROC 0.651.
  • MMLU accuracy / ECE, logprobs (direct route)83.4% / 0.026held outn = 1,000AUROC 0.863. Samples k=4: ECE 0.053, AUROC 0.768.
  • Hosted route, samples k=4: BoolQ / MMLU accuracy90.5% / 84.0%held outn = 200200 questions each; ECE 0.032 / 0.028; $0.61 / $0.58 per 1,000.
  • SummEval rubric, mean Spearman vs expert mean (logprobs)0.525held outn = 2525 held-out articles, 1,600 judgments.
  • MT-Bench pairwise agreement with expert votes, with ties (logprobs)64.8%held outn = 60081.3% without ties; GPT-4 pair judge on the same rows 65.2%. Pairwise probabilities are poorly calibrated (ECE 0.17).
  • Same answer when the temperature-0 call is repeated, BoolQ / MMLU (direct route)99.0% / 95.7%held outn = 1,000

Dataset

Public benchmarks with fixed dev/test splits (seed 24): BoolQ (300 dev, 1,000 test), MMLU (300 dev, 1,000 test), SummEval (10 dev, 25 test articles), MT-Bench human judgments (150 dev, 600 test votes). Calibration fit on dev only; each test split run once.

Caveats

  • Logprobs, the only method that ranks right against wrong answers well, need the direct route (self-host) today; the hosted API uses samples.
  • Stated confidence is almost always 95-100 and carries little information.
  • A busy inference server is not bit-reproducible: repeated temperature-0 calls change some answers, mostly near ties.
  • Only these public benchmarks: calibration varies by domain, so check on your own labelled data. Multi-label use was not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
4.1 s
Receipts
35
Model calls
n/a
Tokens
n/a
Cost per run
$0.004

Self-host verification

Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers

Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. "method": "auto" used logprobs as documented (7 calls for the ticket); two runs gave the same verdicts hash, probabilities moved by up to 0.02; the record verifies as signed by this box and a changed answer fails.

Rehearsal bundle: typed-judgment.zip (2 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Speed depends on load: the support-ticket sample (4 questions, 4 samples each, 35 calls) took about 4 s on a quiet GPU and 25-80 s while the shared GPU was busy (25 Sep 2026).
  • Log-probabilities, the best-calibrated method, are self-host only until the gateway passes them through; hosted requests use samples or stated confidence.
  • Calibration was measured on public benchmarks and varies by domain: check it on your own labelled data before a threshold decides anything.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)decosa-api judgment module (decosa_api/verticals/judgment)AGPL-3.0-or-later
  • Judge: one temperature-0 call per question, plus seeded samplesQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one call per question (stated confidence) (2)
  • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped): 90.7% / 0.076 / 0.088 / 0.703docs/evals/typed-judgment.md, calibration fit on separate dev splits
  • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped): 83.4% / 0.124 / 0.152 / 0.602docs/evals/typed-judgment.md, calibration fit on separate dev splits
Standard · hosted, answer plus 4 seeded samples (5)
  • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples): 90.7% / 0.023 / 0.079 / 0.651docs/evals/typed-judgment.md, calibration fit on separate dev splits
  • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples): 83.4% / 0.053 / 0.114 / 0.768docs/evals/typed-judgment.md, calibration fit on separate dev splits
  • BoolQ, 200 questions through the hosted gateway route: accuracy / ECE / Brier: 90.5% / 0.032 / 0.077docs/evals/typed-judgment.md (service run)
  • MMLU, 200 questions through the hosted gateway route: accuracy / ECE / Brier: 84.0% / 0.028 / 0.121docs/evals/typed-judgment.md (service run)
  • MT-Bench pairwise, samples k=4: agreement with ties / without ties: 64.5% / 81.0%docs/evals/typed-judgment.md
Best · self-host with logprobs (4)
  • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled): 90.7% / 0.023 / 0.069 / 0.857docs/evals/typed-judgment.md, calibration fit on separate dev splits
  • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled): 83.4% / 0.026 / 0.102 / 0.863docs/evals/typed-judgment.md, calibration fit on separate dev splits
  • SummEval rubric (25 held-out articles): mean per-article Spearman with experts, expected score: 0.525 (logprobs) · 0.491 (samples, hosted)docs/evals/typed-judgment.md; G-Eval with GPT-4 reported 0.514 on the full set (different protocol)
  • MT-Bench pairwise (600 held-out expert votes): agreement with ties / without ties: 64.8% / 81.3%; GPT-4 judge on the same rows 65.2% / 83.6%docs/evals/typed-judgment.md; GPT-4 verdicts from the dataset's gpt4_pair split
Wanted · two large judges that must agree (1)
  • BoolQ and MMLU accuracy and calibration, same splits as standard: not measured yet

How we measure · All tools