{"schema_version":"1","site":"https://decosa.ai","id":"typed-judgment","num":"24","name":"Typed-judgment API","tool_name":"Classify with a confidence you can act on","short":"Typed judgments","blurb":"Yes/no, one-of-N, score and label questions that come back as typed answers with a calibrated probability, a short reason and a receipt. It runs on pinned open weights, so a request can be re-run next quarter against the same model, and confident verdicts reproduce. An eval mode grades an output against a rubric or compares two.","status":"live","labels":{"industry":["general","software"],"job":["review","attest"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["data","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["general","software"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["typed-judgment","signed-record"],"models":"Qwen3.8-27B","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the API runs on CPU","final_artifact":"Typed answers with probabilities and a signed record you can re-run and compare.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; signed record with the weights, prompt and seed","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":4126,"p95_ms":null,"runs":null,"receipts_per_run":35,"cost_per_run_usd":0.0044},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. \"method\": \"auto\" used logprobs as documented (7 calls for the ticket); two runs gave the same verdicts hash, probabilities moved by up to 0.02; the record verifies as signed by this box and a changed answer fails."},"known_limits":["Speed depends on load: the support-ticket sample (4 questions, 4 samples each, 35 calls) took about 4 s on a quiet GPU and 25-80 s while the shared GPU was busy (25 Sep 2026).","Log-probabilities, the best-calibrated method, are self-host only until the gateway passes them through; hosted requests use samples or stated confidence.","Calibration was measured on public benchmarks and varies by domain: check it on your own labelled data before a threshold decides anything."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"BoolQ accuracy / ECE, logprobs (direct route)","value":"90.7% / 0.023","unit":null,"n":1000,"split":"heldout","note":"AUROC 0.857. Samples k=4 (hosted default): ECE 0.023, AUROC 0.651."},{"name":"MMLU accuracy / ECE, logprobs (direct route)","value":"83.4% / 0.026","unit":null,"n":1000,"split":"heldout","note":"AUROC 0.863. Samples k=4: ECE 0.053, AUROC 0.768."},{"name":"Hosted route, samples k=4: BoolQ / MMLU accuracy","value":"90.5% / 84.0%","unit":null,"n":200,"split":"heldout","note":"200 questions each; ECE 0.032 / 0.028; $0.61 / $0.58 per 1,000."},{"name":"SummEval rubric, mean Spearman vs expert mean (logprobs)","value":"0.525","unit":null,"n":25,"split":"heldout","note":"25 held-out articles, 1,600 judgments."},{"name":"MT-Bench pairwise agreement with expert votes, with ties (logprobs)","value":"64.8%","unit":null,"n":600,"split":"heldout","note":"81.3% without ties; GPT-4 pair judge on the same rows 65.2%. Pairwise probabilities are poorly calibrated (ECE 0.17)."},{"name":"Same answer when the temperature-0 call is repeated, BoolQ / MMLU (direct route)","value":"99.0% / 95.7%","unit":null,"n":1000,"split":"heldout","note":null}],"dataset":"Public benchmarks with fixed dev/test splits (seed 24): BoolQ (300 dev, 1,000 test), MMLU (300 dev, 1,000 test), SummEval (10 dev, 25 test articles), MT-Bench human judgments (150 dev, 600 test votes). Calibration fit on dev only; each test split run once.","held_out":true,"caveats":["Logprobs, the only method that ranks right against wrong answers well, need the direct route (self-host) today; the hosted API uses samples.","Stated confidence is almost always 95-100 and carries little information.","A busy inference server is not bit-reproducible: repeated temperature-0 calls change some answers, mostly near ties.","Only these public benchmarks: calibration varies by domain, so check on your own labelled data. Multi-label use was not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/typed-judgment"},"stack":{"summary":"Send a context and up to 16 narrow questions. Each comes back as a typed answer (yes or no, an option key, a level, a set of labels) with a probability, the full distribution, a one-line reason and the receipt of every model call, and the request ends with a signed record of the weights, prompt, seed and answers. Because the weights and the prompt are pinned and recorded, the same request can be re-run later against the same model, which a closed judge that is updated or retired cannot offer; in our re-runs no answer with a probability of 0.9 or more changed, and near-ties are flagged. An eval mode grades an output against a rubric or compares two outputs in both orders. The probabilities were calibrated on public dev sets and measured on held-out test sets; the numbers are below.","tagline":"Yes/no, one-of-N, score and label questions answered with a calibrated probability, a reason and a receipt, on pinned open weights.","deployment":"hosted-or-self-host","regulatory_note":"The answers are model outputs with estimated probabilities, not facts or decisions: keep the decision rule and any human review in your own code. Calibration was measured on public benchmarks (BoolQ, MMLU, SummEval, MT-Bench) and varies by domain, so check it on your own labelled data before a threshold decides anything that matters. Where a judgment feeds a decision about a person (hiring, credit, housing, insurance), rules such as NYC Local Law 144 (in force), Illinois HB 3773 (in force 1 Jan 2026), Colorado's AI Act (effective date moved to 1 Jan 2027) and the EU AI Act's high-risk duties (Annex III uses from 2 Dec 2027) can apply to you as the deployer; the signed records and receipts help with record-keeping but are not a compliance programme. The hosted demo keeps no context: it stays in memory for the request, the record holds hashes only, and logs carry counts. For personal or confidential data, self-host. Not legal advice. Model licence: Apache-2.0 (Qwen3.8-27B). Dataset licences for the eval: BoolQ CC BY-SA 3.0, MMLU MIT, SummEval MIT, MT-Bench human judgments CC BY 4.0. Checked 25 Sep 2026.","components":[{"id":"engine","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","name":"decosa-api judgment module (decosa_api/verticals/judgment)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","best","wanted"],"alternative_to":null},{"id":"judge","role":"Judge: one temperature-0 call per question, plus seeded samples","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, thinking off, prefix caching, top-20 logprobs","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard","best"],"alternative_to":null},{"id":"gemma-judge","role":"Fast judge with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0","params":"26B","quant":"BF16 (49 GB)","vram_gb":49,"memory_gb_estimate":null,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null},{"id":"dsv41-wanted","role":"Second judge for disagreement","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"glm-wanted","role":"Third judge, from another family","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one call per question (stated confidence)","summary":"The cheapest hosted setting (method single): the same answers, but the probability is the model's stated confidence mapped on a dev set, which separates right from wrong answers poorly.","components":["engine","judge"],"hardware":"Hosted; or 1x RTX 5090 32 GB (estimate) / 1x RTX PRO 6000 96 GB (measured) self-hosted","quality_evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)","value":"90.7% / 0.076 / 0.088 / 0.703","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)","value":"83.4% / 0.124 / 0.152 / 0.602","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"}],"latency_note":"measured: one call per question, like the best tier.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"One gateway-signed receipt per question.","hosting":null},{"id":"standard","label":"Standard · hosted, answer plus 4 seeded samples","summary":"What the hosted API runs by default (method samples, k=4): five calls per question, probabilities from smoothed votes. Better calibrated than lite, coarser than logprobs.","components":["engine","judge"],"hardware":"Hosted (1x RTX PRO 6000 96 GB behind the gateway)","quality_evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)","value":"90.7% / 0.023 / 0.079 / 0.651","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)","value":"83.4% / 0.053 / 0.114 / 0.768","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"BoolQ, 200 questions through the hosted gateway route: accuracy / ECE / Brier","value":"90.5% / 0.032 / 0.077","source":"docs/evals/typed-judgment.md (service run)"},{"metric":"MMLU, 200 questions through the hosted gateway route: accuracy / ECE / Brier","value":"84.0% / 0.028 / 0.121","source":"docs/evals/typed-judgment.md (service run)"},{"metric":"MT-Bench pairwise, samples k=4: agreement with ties / without ties","value":"64.5% / 81.0%","source":"docs/evals/typed-judgment.md"}],"latency_note":"measured: the samples run in parallel with the main call; with prefix caching they add little time on a quiet card.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every call, including each sample, has its own gateway-signed receipt.","hosting":null},{"id":"best","label":"Best · self-host with logprobs","summary":"On your own box the engine reads the answer token's log-probabilities: one call per question and the best-calibrated probabilities in our eval. Hosted too once the gateway passes logprobs through.","components":["engine","judge"],"hardware":"1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)","value":"90.7% / 0.023 / 0.069 / 0.857","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)","value":"83.4% / 0.026 / 0.102 / 0.863","source":"docs/evals/typed-judgment.md, calibration fit on separate dev splits"},{"metric":"SummEval rubric (25 held-out articles): mean per-article Spearman with experts, expected score","value":"0.525 (logprobs) · 0.491 (samples, hosted)","source":"docs/evals/typed-judgment.md; G-Eval with GPT-4 reported 0.514 on the full set (different protocol)"},{"metric":"MT-Bench pairwise (600 held-out expert votes): agreement with ties / without ties","value":"64.8% / 81.3%; GPT-4 judge on the same rows 65.2% / 83.6%","source":"docs/evals/typed-judgment.md; GPT-4 verdicts from the dataset's gpt4_pair split"}],"latency_note":"measured: one call per question.","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Self-hosted calls are signed by your own box (attested), not countersigned by the gateway.","hosting":null},{"id":"wanted","label":"Wanted · two large judges that must agree","summary":"DeepSeek-V4.1-Flash and GLM-5.3-Flash answer beside Qwen3.8-27B; when the families disagree the answer is flagged instead of returned with one model's probability. Not served yet.","components":["engine","judge","dsv41-wanted","glm-wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.","quality_evidence":[{"metric":"BoolQ and MMLU accuracy and calibration, same splits as standard","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[{"id":"fast-judge","label":"Fast option-token judge","components":["gemma-judge"],"hardware":"1x RTX PRO 6000 96 GB (BF16 weights are 49 GB)","use":"A 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway.","status":"not served"}],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /judgment/info, /judgment/samples; POST /judgment/judge (SSE or JSON), /judgment/eval, /judgment/batch, /judgment/v1/systemone, /judgment/verify. Keeps no context."},{"name":"vLLM (judge)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly with logprobs (self-host)."}],"tools":[{"name":"BoolQ (google/boolq)","url":"https://huggingface.co/datasets/google/boolq","license":"CC BY-SA 3.0","purpose":"Yes/no eval: 300 validation questions to fit calibration, 1,000 others held out."},{"name":"MMLU (cais/mmlu)","url":"https://huggingface.co/datasets/cais/mmlu","license":"MIT","purpose":"One-of-four eval: 300 validation questions to fit calibration, 1,000 test questions held out."},{"name":"SummEval (mteb/summeval)","url":"https://huggingface.co/datasets/mteb/summeval","license":"MIT","purpose":"Eval mode, rubric: expert 1-5 ratings of 16 summaries per article on four criteria. 10 articles for dev, 25 held out."},{"name":"MT-Bench human judgments (lmsys/mt_bench_human_judgments)","url":"https://huggingface.co/datasets/lmsys/mt_bench_human_judgments","license":"CC BY 4.0","purpose":"Eval mode, pairwise: expert votes between two chat answers, compared with our judge and with the GPT-4 judge on the same rows. 600 votes held out."},{"name":"scripts/judgment_eval.py and docs/evals/typed-judgment.md","url":null,"license":"Apache-2.0","purpose":"Rebuilds the splits, runs the model, fits calibration on dev only, and writes the metrics, the repeat-agreement checks and the calibration file."},{"name":"POST /judgment/verify","url":null,"license":"Apache-2.0","purpose":"Checks a record's signature against this server's key and, if you send it, the context hash. The console also checks the signature in your browser with WebCrypto."}],"hardware":[{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code tool, not run here for this tool."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services."}],"latency":[{"lane":"one question, one call, direct route, card shared","typical_ms":3537,"source":"measured on our server 2026-09-25: median of the BoolQ test calls (about 270 prompt tokens by vLLM's count), 16 in flight, other evals on the same card"},{"lane":"support-ticket sample (4 questions + 4 labels, samples k=4: 35 calls), hosted gateway route","typical_ms":35000,"source":"measured on our server 2026-09-25 while the shared card had 30-70 requests queued; the calls of one request run 8 at a time, so a quiet card is several times faster"}],"benchmark":null,"notes":["Logprobs are clearly the best probability: on held-out BoolQ and MMLU they rank right against wrong answers far better (AUROC 0.857 and 0.863) than 4 samples (0.651 and 0.768) or the model's stated confidence (0.703 and 0.602), which is almost always 95 to 100. Samples are well calibrated on average (ECE 0.023 on BoolQ) but coarse. That gap is the case for the gateway logprobs passthrough.","Re-runs are not bit-identical on a shared, busy server: batching changes the arithmetic. Repeating the temperature-0 call on 1,000 BoolQ and 1,000 MMLU test questions kept 99.0% and 95.7% of answers; every change was on an answer the calibrated logprobs put under 0.85, and none of the 1,274 answers at 0.9 or more changed. Through the hosted route (answer plus 4 seeded samples, 400 questions, run twice) 98.5% and 97.5% of answers held, 93% and 83% of distributions were identical, and none of the 182 answers at 0.9 or more changed. The API marks answers under 0.75 as near_tie. Pinned weights and prompts rule out the other kind of drift: a closed model being replaced.","Measured through the gateway at its list price for Qwen3.8-27B ($0.30 per million prompt tokens, $1.50 per million generated): a yes/no or choice question with a ~300-token context costs about $0.12 per 1,000 judgments with one call and no reason, about $0.17 with a reason, and $0.58-0.61 per 1,000 with the default 4 samples (5 calls, each paying the prompt). Self-hosted with logprobs: one call, GPU time only. Batches of up to 50 items report the exact cost per item.","Eval mode on held-out data: on 25 SummEval articles (400 summaries, four criteria) the expected 1-5 score has a mean per-article Spearman correlation of 0.525 with expert ratings (logprobs; 0.491 with samples). On 600 MT-Bench expert votes the pairwise judge agreed 64.8% with ties and 81.3% without ties; the GPT-4 judge on the same rows agreed 65.2% and 83.6%. Both orders picked the same winner 92% of the time.","Calibration was fit on the dev splits only (a logprob temperature, a vote prior and a stated-confidence table per answer type); the test splits were run once. Scores (1-5) use a temperature and prior chosen on SummEval dev articles; labels use the yes/no settings and are not separately measured.","The gateway strips logprobs today. docs/proposals/gateway-logprobs.patch in decosa-api is a proposed 40-line change to the gateway that forwards them on non-streaming calls; the existing gateway tests pass with it, but it is not applied."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: contexts and questions live in memory for the request. The signed record holds hashes of the context and questions plus the answers, not your text."},{"label":"What leaves the box","value":"Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box."},{"label":"Input formats","value":"JSON: a context (text or any JSON, up to 24,000 characters) and 1-16 typed questions (yes/no, one of up to 26 options, a 2-9 level score, or up to 12 labels); or a rubric or a pair of outputs to grade; batches of up to 50 items."},{"label":"Typical run","value":"The support ticket sample: a few dozen calls and a fraction of a cent with repeated samples, or a handful of calls and less with stated confidence, at the gateway list price."}],"data_handling":{"page":"/data#typed-judgment","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: contexts and questions live in memory for the request. The signed record holds hashes of the context and questions plus the answers, not your text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/developer/typed-judgment","input":"judgment","lanes":[{"id":"answers","title":"Typed answers","kind":"list"},{"id":"record","title":"Signed record","kind":"json"},{"id":"rerun","title":"Run again","kind":"markdown"}],"samples":[{"n":1,"id":"support-triage","title":"Support triage","deep_link":"/tools/developer/typed-judgment?sample=1&autorun=0"},{"n":2,"id":"agent-done-check","title":"Agent done check","deep_link":"/tools/developer/typed-judgment?sample=2&autorun=0"},{"n":3,"id":"lead-fit","title":"Lead fit","deep_link":"/tools/developer/typed-judgment?sample=3&autorun=0"},{"n":4,"id":"eval-rubric","title":"Eval rubric","deep_link":"/tools/developer/typed-judgment?sample=4&autorun=0"},{"n":5,"id":"eval-pairwise","title":"Eval pairwise","deep_link":"/tools/developer/typed-judgment?sample=5&autorun=0"},{"n":6,"id":"moderation","title":"Moderation","deep_link":"/tools/developer/typed-judgment?sample=6&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/typed-judgment-hosted.md","selfhost":"/prompts/typed-judgment-selfhost.md","assemble":"/prompts/typed-judgment-assemble.md","mac":"/prompts/typed-judgment-mac.md"},"rehearsal":{"bundle":"/samples/typed-judgment.zip","bundle_url":"https://decosa.ai/samples/typed-judgment.zip","folder":"/samples/typed-judgment/","expected":"/samples/typed-judgment/expected.json","files":["/samples/typed-judgment/expected.json","/samples/typed-judgment/inputs/questions.json","/samples/typed-judgment/inputs/ticket.txt"],"bytes":2204,"checks":["the customer asks for a refund: yes","the refund answer carries a probability of at least 0.8","the returns team handles it first","urgency is 'this week' or 'within two days'","the tags include damaged and double-charge","every tag has its own probability","the signed record verifies","the record matches the ticket","a record with the refund answer changed no longer verifies","every model call has a signed receipt"],"licence":"Fictional: the customer, the shop and the order are made up for Decosa. Part of decosa-api, AGPL-3.0-or-later.","about":"A fictional customer's support ticket: a kettle arrived cracked, they want their money back before Friday, and delivery was charged twice. Four typed questions (yes/no, choice, score, labels) must come back with answers and probabilities (refund yes, the returns team, urgency this week or sooner, tags damaged and double-charge), and the signed record must verify against the ticket and catch a changed answer.","run":{"containers":"docker compose exec api python scripts/rehearse.py typed-judgment","checkout":"python scripts/rehearse.py typed-judgment --bundle typed-judgment.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py typed-judgment"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=typed-judgment","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1377.6,"basis":"estimate","unknown":[]},{"id":"alternate-fast-judge","gpu_gb":57,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/typed-judgment","json":"/use-cases/typed-judgment.json","metrics":"/metrics/typed-judgment","console":"/tools/developer/typed-judgment","console_sample":"/tools/developer/typed-judgment?sample=1&autorun=0","stack":"/tools/developer/typed-judgment#stack","try_live":"/tools/developer/typed-judgment","watch":"/tools/developer/typed-judgment","build":"/tools/developer/typed-judgment#build","self_host":"/tools/developer/typed-judgment#self-host","prompts":{"hosted":"/prompts/typed-judgment-hosted.md","selfhost":"/prompts/typed-judgment-selfhost.md","assemble":"/prompts/typed-judgment-assemble.md","mac":"/prompts/typed-judgment-mac.md"}}}