Skip to content
decosa
LiveHostedSelf-hostMac

Classify with a confidence you can act on

Yes/no, one-of-N, score or label answers with a probability you can set a threshold on and a reason, from a pinned model you can re-run next quarter.

Measured90.7% / 0.023Yes/no questions (BoolQ): accuracy / calibration error
On production4.1 smedian on production (2026-09-25); slower when the service is busy
List price~$0.60 per 1,000 questionsmeasured, at list price

Built on: Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the typed-judgment api API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the API runs on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
typed-judgment

Use the hosted API

# Decosa Typed-judgment API: use the hosted API

You are wiring Decosa's typed-judgment API into this project. It answers narrow questions about a context with typed
answers: `yesno`, `choice` (one of N options), `score` (ordered levels, 1 to 5 by default) or `labels` (any of N). Each
answer comes with a probability, the full distribution, a short reason and receipt ids, and each request ends with a
record signed by the server. The model is Qwen3.8-27B on pinned open weights, so the same request with the same seed can be
re-run later against the same model (confident verdicts reproduce; see the limits). An eval mode grades an output against a rubric, or compares two outputs. Use only what is listed
below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- Keep policy in code: the API returns probabilities; thresholds, weights and "what happens next" belong to this project.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "typed-judgment"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`). Over a limit: HTTP 429 with `Retry-After`.
3. Budget per request: about 90 generated tokens per question (16 with `"reasons": false`) plus 8 per sample. 402 if it
   would not fit. One request at a time per demo token; a key allows 2 at once and 60 requests a minute.

## Endpoints
- `POST /judgment/judge` (token). Body:
  `{"context"?: "text or any JSON", "questions": {"<id>": {"type": "yesno"|"choice"|"score"|"labels", "question": "...", "options"?: {"key": "description"}, "levels"?: ["lowest", "...", "highest"] | 5, "labels"?: {"key": "description"}, "threshold"?: 0.5}}, "method"?: "auto"|"samples"|"single"|"logprobs", "samples"?: 2-8, "seed"?: 0, "reasons"?: true}`
  - Limits: 16 questions, 26 options, 2-9 levels, 12 labels (each label is one model call), context 24,000 characters.
  - `method`: `samples` (the hosted default: the temperature-0 answer plus k seeded samples, 4 by default), `single` (one
    call, the model's stated confidence mapped through a table fit on a dev set), `logprobs` (self-host only until the
    gateway passes log-probabilities through; 400 on the hosted API).
  - JSON by default: `{answers: {<id>: {type, answer, probability, distribution, reason, method, near_tie, receipt_ids, expected? (score), option? (choice), probabilities? (labels)}}, record, usage: {calls, prompt_tokens, completion_tokens, cost_usd}, receipts, budget, note}`.
    A question that could not be judged has `answer: null` and an `error`; the others still come back.
  - With `Accept: text/event-stream` (or `"stream": true`): `ready`, then a `receipt` per model call and an `answer` per
    question as each finishes, then `result` (answers, record, usage), `budget`, `done`.
- `POST /judgment/eval` (token), eval-judge mode, same output shape:
  - rubric: `{"task"?: "...", "reference"?: "...", "candidate": "...", "rubric": [{"id": "faithful", "question": "...", "type"?: "score"}]}`
    Criteria are 1-5 scores (very poor to excellent) unless they say otherwise.
  - pairwise: `{"task"?: "...", "candidates": ["answer A", "answer B"]}` → `answers.preference` with `answer` `a`, `b` or
    `tie`, the averaged distribution over both orders, and `position_consistent`.
- `POST /judgment/batch` (token): `{"items": [{"id": "row-1", "context": ..., "questions": {...}}], "method"?, "samples"?, "seed"?, "reasons"?}`.
  Up to 50 items and 400 model calls per batch. JSON `{results: [{id, answers, record, usage}], usage: {items, calls, cost_usd}}`,
  or NDJSON lines (`item` ..., `usage`, `budget`, `done`) with `"stream": true`.
- `POST /judgment/v1/systemone` (token): the local decision server's shape (`{state, questions: {id: {type: noul|choice|score, instructions, criteria}}, options: {seed, samples}}`),
  answered in that shape (`noul`, `choice`, `score`, `probabilities`, `confidence`) plus `record`. Score criteria are zero-based there.
- `POST /judgment/verify` (no token) `{"record": {...}, "context"?: ...}` → `{valid_signature, signed_by_this_server, context_matches?}`.
- `GET /judgment/info` (methods, limits, prices, the calibration in use and the held-out numbers), `GET /judgment/samples`,
  `GET /attest/signing-key` (no token).

## Example: route support tickets (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {
    "context": ticket_text,
    "questions": {
        "refund": {"type": "yesno", "question": "Does the customer ask for their money back?"},
        "team": {"type": "choice", "question": "Which team should handle it?",
                 "options": {"returns": "damaged or unwanted items", "billing": "charges", "shipping": "late or lost parcels"}},
    },
    "method": "samples", "samples": 4, "seed": 0,
}
r = httpx.post(f"{API}/judgment/judge", json=body, headers=H, timeout=120)
r.raise_for_status()
out = r.json()
team = out["answers"]["team"]
if team["probability"] < 0.7:
    send_to_human(ticket_id, reason=team["reason"])          # the threshold is yours
else:
    route(ticket_id, team["answer"])
store(ticket_id, out["record"])                                # keep the signed record with the decision
```

## Same thing in TypeScript (Node 18+ fetch)
```ts
const res = await fetch("https://api.decosa.ai/judgment/judge", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.DECOSA_API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({ context: ticket, questions: { refund: { type: "yesno", question: "Does the customer ask for their money back?" } } }),
});
if (!res.ok) throw new Error(`${res.status}: ${(await res.json()).error}`);
const { answers, record, usage } = await res.json();
console.log(answers.refund.answer, answers.refund.probability, record.verdicts_sha256, usage.cost_usd);
```

## And curl
```bash
curl -sS https://api.decosa.ai/judgment/judge -H "Authorization: Bearer $DECOSA_API_KEY" -H "Content-Type: application/json" \
  -d '{"context":"The parcel arrived two weeks late.","questions":{"late":{"type":"yesno","question":"Was the delivery late?"}}}'
```

## Re-run and compare (the reproducibility check)
Send the same request with the same `seed` again and compare `record.verdicts_sha256` (answers only) and
`record.answers_sha256` (answers and probabilities). Also check `record.method.model_root` (the weights) and
`record.method.prompt_sha256`: if either changed, the model or the prompt changed, and scores are not comparable.

## Verify a record yourself (`pip install cryptography`)
```python
import json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
record = out["record"]                                         # from the request above
pub = json.load(urllib.request.urlopen("https://api.decosa.ai/attest/signing-key"))["pubkey"]
assert record["signer"] == pub
body = {k: v for k, v in record.items() if k != "sig"}
msg = record["v"].encode() + b"\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(record["sig"]), msg)
```
Probabilities in the record are integers in basis points (10000 = 1.0). Each id in `record["receipts"]` resolves at
`GET https://api.decosa.ai/receipts/{id}` to the gateway-signed receipt for that model call.

## Honest limits
- Re-runs are not bit-identical on a busy shared server (batching changes the arithmetic). In our eval, 98-99% of yes/no
  answers and 96-98% of four-option answers held on a re-run, and no answer with probability 0.9 or more changed. Answers
  under 0.75 come back with `near_tie: true` (labels: a list); treat them as uncertain, not as verdicts.
- Cost at the gateway's list price, measured: about $0.12 per 1,000 one-call judgments without a reason, $0.17 with one,
  and about $0.60 per 1,000 with the default 4 samples. `usage.cost_usd` gives the exact figure per request.
- Probabilities were calibrated on public dev sets (BoolQ for yes/no, MMLU for choices) and measured on held-out test
  sets: see the Stack tab. Scores (1-5) and labels use the yes/no and default settings and are not separately calibrated.
  Calibration varies by domain: log your own outcomes and check before a threshold decides anything that matters.
- With `samples`, probabilities are coarse (k+1 votes, smoothed). More samples cost more prompt tokens.
- The model trusts the context you send. Put facts in the context; do not ask it to know things it cannot see.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Typed-judgment API: run it yourself (containers)

You are setting up the Decosa typed-judgment API on this machine, so contexts never leave it and the model weights stay
pinned. It answers typed questions (yes/no, one of N, scores, labels) with calibrated probabilities and signs a record
per request. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/typed-judgment.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py typed-judgment` (the api image carries the same bundle under /app/rehearsal/typed-judgment/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py typed-judgment --bundle typed-judgment.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer asks for a refund: yes", "the refund answer carries a probability of at least 0.8", "the returns team handles it first"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, weights revision `482ca0f3832238542f8f5295dde86b5f22711d80`)
   and the `api` service. For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to 127.0.0.1. Keep `--enable-prefix-caching` on the `llm` service.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/judgment/info` must show `methods.auto: "logprobs"` (the direct route reads
   token log-probabilities: one call per question, the best-calibrated method). `GET /attest/signing-key` shows this box's
   public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"typed-judgment"}`, fetch `GET /judgment/samples`, and
   send the `support-triage` sample's `request` (without `method` and `samples`) to `POST /judgment/judge`. Expect
   `refund: yes` and `team: returns`. Send it again and compare `record.verdicts_sha256`: it must match. Then
   `POST /judgment/verify` with the record: `valid_signature` and `signed_by_this_server` must both be true.
6. Report back: the public key and key id, the answers with their probabilities, whether the re-run matched, and how
   long each request took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential contexts. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/typed-judgment-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090best tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.) The best tier fits too.

  • GeForce RTX 5090best tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too.

  • 2x GeForce RTX 5090best tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size). The best tier fits too.

  • L40Sbest tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.

  • H100 80 GB (SXM)best tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.

  • RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 96 GB). The best tier fits too.

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBbest tierRuns

    The standard tier fits (32 of 96 GB). The best tier fits too.

  • Apple M5 Max, 64 GBbest tierRuns

    The standard tier fits (32 of 64 GB). The best tier fits too.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"typed-judgment"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py typed-judgment

Download the mock-data bundle (2 KB, 10 checks)expected.json

A fictional customer's support ticket: a kettle arrived cracked, they want their money back before Friday, and delivery was charged twice. Four typed questions (yes/no, choice, score, labels) must come back with answers and probabilities (refund yes, the returns team, urgency this week or sooner, tags damaged and double-charge), and the signed record must verify against the ticket and catch a changed answer.

What the rehearsal checks
  • the customer asks for a refund: yes
  • the refund answer carries a probability of at least 0.8
  • the returns team handles it first
  • urgency is 'this week' or 'within two days'
  • the tags include damaged and double-charge
  • every tag has its own probability
  • the signed record verifies
  • the record matches the ticket
  • a record with the refund answer changed no longer verifies
  • every model call has a signed receipt

Licence: Fictional: the customer, the shop and the order are made up for Decosa. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Typed-judgment API: run it yourself (containers)

You are setting up the Decosa typed-judgment API on this machine, so contexts never leave it and the model weights stay
pinned. It answers typed questions (yes/no, one of N, scores, labels) with calibrated probabilities and signs a record
per request. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/typed-judgment.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py typed-judgment` (the api image carries the same bundle under /app/rehearsal/typed-judgment/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py typed-judgment --bundle typed-judgment.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer asks for a refund: yes", "the refund answer carries a probability of at least 0.8", "the returns team handles it first"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
   check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, weights revision `482ca0f3832238542f8f5295dde86b5f22711d80`)
   and the `api` service. For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
   `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to 127.0.0.1. Keep `--enable-prefix-caching` on the `llm` service.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/judgment/info` must show `methods.auto: "logprobs"` (the direct route reads
   token log-probabilities: one call per question, the best-calibrated method). `GET /attest/signing-key` shows this box's
   public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"typed-judgment"}`, fetch `GET /judgment/samples`, and
   send the `support-triage` sample's `request` (without `method` and `samples`) to `POST /judgment/judge`. Expect
   `refund: yes` and `team: returns`. Send it again and compare `record.verdicts_sha256`: it must match. Then
   `POST /judgment/verify` with the record: `valid_signature` and `signed_by_this_server` must both be true.
6. Report back: the public key and key id, the answers with their probabilities, whether the re-run matched, and how
   long each request took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential contexts. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/typed-judgment-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsTyped-judgment API on GeForce RTX 5090: use the Best · self-host with logprobs tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code use case, not run here for this use case.

Best · self-host with logprobs: what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Engine: decosa-api judgment module (decosa_api/verticals/judgment). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Typed-judgment API, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Typed-judgment API on my hardware

Fetch https://decosa.ai/prompts/typed-judgment-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=typed-judgment)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Best · self-host with logprobs (best). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Engine: decosa-api judgment module (decosa_api/verticals/judgment), CPU
- Judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/typed-judgment-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Typed-judgment API: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Typed-judgment API on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/typed-judgment.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py typed-judgment` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer asks for a refund: yes", "the refund answer carries a probability of at least 0.8", "the returns team handles it first"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Judge: one temperature-0 call per question, plus seeded samples | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py typed-judgment`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Logprobs work with mlx_lm.server (10 alternatives per token). With oMLX, which returns none, judgments use the samples method.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 4.1 s · ~$0.004 per run · 35 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone, compose up, sample against local model servers

Measured cost to run: about $0.60 per 1,000 questions (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. "method": "auto" used logprobs as documented (7 calls for the ticket); two runs gave the same verdicts hash, probabilities moved by up to 0.02; the record verifies as signed by this box and a changed answer fails.

Known limits (3)
  • Speed depends on load: the support-ticket sample (4 questions, 4 samples each, 35 calls) took about 4 s on a quiet GPU and 25-80 s while the shared GPU was busy (25 Sep 2026).
  • Log-probabilities, the best-calibrated method, are self-host only until the gateway passes them through; hosted requests use samples or stated confidence.
  • Calibration was measured on public benchmarks and varies by domain: check it on your own labelled data before a threshold decides anything.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the typed-judgment api API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the API runs on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Yes/no, one-of-N, score and label questions answered with a calibrated probability, a reason and a receipt, on pinned open weights.

Send a context and up to 16 narrow questions. Each comes back as a typed answer (yes or no, an option key, a level, a set of labels) with a probability, the full distribution, a one-line reason and the receipt of every model call, and the request ends with a signed record of the weights, prompt, seed and answers. Because the weights and the prompt are pinned and recorded, the same request can be re-run later against the same model, which a closed judge that is updated or retired cannot offer; in our re-runs no answer with a probability of 0.9 or more changed, and near-ties are flagged. An eval mode grades an output against a rubric or compares two outputs in both orders. The probabilities were calibrated on public dev sets and measured on held-out test sets; the numbers are below.

Deployment
Hosted or self-host
Regulatory
The answers are model outputs with estimated probabilities, not facts or decisions: keep the decision rule and any human review in your own code. Calibration was measured on public benchmarks (BoolQ, MMLU, SummEval, MT-Bench) and varies by domain, so check it on your own labelled data before a threshold decides anything that matters. Where a judgment feeds a decision about a person (hiring, credit, housing, insurance), rules such as NYC Local Law 144 (in force), Illinois HB 3773 (in force 1 Jan 2026), Colorado's AI Act (effective date moved to 1 Jan 2027) and the EU AI Act's high-risk duties (Annex III uses from 2 Dec 2027) can apply to you as the deployer; the signed records and receipts help with record-keeping but are not a compliance programme. The hosted demo keeps no context: it stays in memory for the request, the record holds hashes only, and logs carry counts. For personal or confidential data, self-host. Not legal advice. Model licence: Apache-2.0 (Qwen3.8-27B). Dataset licences for the eval: BoolQ CC BY-SA 3.0, MMLU MIT, SummEval MIT, MT-Bench human judgments CC BY 4.0. Checked 25 Sep 2026.
Architecture
Text description

A context with typed questions, a rubric with an output to grade, two outputs to compare, or a batch of up to 50 items go to the judgment engine. It writes one prompt per question or label and calls Qwen3.8-27B once at temperature 0, plus k seeded samples at temperature 1 when the sampling method is used; on a self-hosted box it reads the answer token's log-probabilities instead. The engine parses each answer, turns votes, stated confidence or log-probabilities into a distribution calibrated on public dev sets, averages pairwise judgments over both orders, and signs a record with the weights root, prompt hash, seed and every answer in basis points. Outputs: typed answers with probabilities and reasons, a record you can compare between re-runs, and one receipt per model call, which our gateway countersigns. In self-host mode the engine and the model run on your machine.

Architecture

At a glance

Data retention
Nothing stored: contexts and questions live in memory for the request. The signed record holds hashes of the context and questions plus the answers, not your text.
What leaves the box
Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box.
Input formats
JSON: a context (text or any JSON, up to 24,000 characters) and 1-16 typed questions (yes/no, one of up to 26 options, a 2-9 level score, or up to 12 labels); or a rubric or a pair of outputs to grade; batches of up to 50 items.
Typical run
The support ticket sample: a few dozen calls and a fraction of a cent with repeated samples, or a handful of calls and less with stated confidence, at the gateway list price.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    one call per question (stated confidence)

    The cheapest hosted setting (method single): the same answers, but the probability is the model's stated confidence mapped on a dev set, which separates right from wrong answers poorly.

    Models
    • decosa-api judgment module (decosa_api/verticals/judgment)
    • Qwen3.8-27B (NVFP4)
    Hardware
    Hosted; or 1x RTX 5090 32 GB (estimate) / 1x RTX PRO 6000 96 GB (measured) self-hosted
    Quality evidence
    • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)90.7% / 0.076 / 0.088 / 0.703docs/evals/typed-judgment.md, calibration fit on separate dev splits
    • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (stated confidence, mapped)83.4% / 0.124 / 0.152 / 0.602docs/evals/typed-judgment.md, calibration fit on separate dev splits
    Latency
    measured: one call per question, like the best tier.
    Verification
    Proof: strongOne gateway-signed receipt per question.
  • In the hosted demo

    Standard

    hosted, answer plus 4 seeded samples

    What the hosted API runs by default (method samples, k=4): five calls per question, probabilities from smoothed votes. Better calibrated than lite, coarser than logprobs.

    Models
    • decosa-api judgment module (decosa_api/verticals/judgment)
    • Qwen3.8-27B (NVFP4)
    Hardware
    Hosted (1x RTX PRO 6000 96 GB behind the gateway)
    Quality evidence
    • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)90.7% / 0.023 / 0.079 / 0.651docs/evals/typed-judgment.md, calibration fit on separate dev splits
    • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (k=4 samples)83.4% / 0.053 / 0.114 / 0.768docs/evals/typed-judgment.md, calibration fit on separate dev splits
    • BoolQ, 200 questions through the hosted gateway route: accuracy / ECE / Brier90.5% / 0.032 / 0.077docs/evals/typed-judgment.md (service run)
    • MMLU, 200 questions through the hosted gateway route: accuracy / ECE / Brier84.0% / 0.028 / 0.121docs/evals/typed-judgment.md (service run)
    • MT-Bench pairwise, samples k=4: agreement with ties / without ties64.5% / 81.0%docs/evals/typed-judgment.md
    Latency
    measured: the samples run in parallel with the main call; with prefix caching they add little time on a quiet card.
    Verification
    Proof: strongEvery call, including each sample, has its own gateway-signed receipt.
  • Best

    self-host with logprobs

    On your own box the engine reads the answer token's log-probabilities: one call per question and the best-calibrated probabilities in our eval. Hosted too once the gateway passes logprobs through.

    Models
    • decosa-api judgment module (decosa_api/verticals/judgment)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • BoolQ yes/no (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)90.7% / 0.023 / 0.069 / 0.857docs/evals/typed-judgment.md, calibration fit on separate dev splits
    • MMLU one of four (1,000 held out): accuracy / ECE / Brier / AUROC (logprobs, temperature-scaled)83.4% / 0.026 / 0.102 / 0.863docs/evals/typed-judgment.md, calibration fit on separate dev splits
    • SummEval rubric (25 held-out articles): mean per-article Spearman with experts, expected score0.525 (logprobs) · 0.491 (samples, hosted)docs/evals/typed-judgment.md; G-Eval with GPT-4 reported 0.514 on the full set (different protocol)
    • MT-Bench pairwise (600 held-out expert votes): agreement with ties / without ties64.8% / 81.3%; GPT-4 judge on the same rows 65.2% / 83.6%docs/evals/typed-judgment.md; GPT-4 verdicts from the dataset's gpt4_pair split
    Latency
    measured: one call per question.
    Verification
    Proof: partialSelf-host onlySelf-hosted calls are signed by your own box (attested), not countersigned by the gateway.
  • Needs more compute

    Wanted: the best setup

    two large judges that must agree

    DeepSeek-V4.1-Flash and GLM-5.3-Flash answer beside Qwen3.8-27B; when the families disagree the answer is flagged instead of returned with one model's probability. Not served yet.

    Models
    • decosa-api judgment module (decosa_api/verticals/judgment)
    • Qwen3.8-27B (NVFP4)
    • DeepSeek-V4.1-Flash
    • GLM-5.3-Flash
    Hardware
    Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.
    Quality evidence
    • BoolQ and MMLU accuracy and calibration, same splits as standardnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • Fast option-token judgeGemma-4-26B-A4B-itnot servedA 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway. Hardware: 1x RTX PRO 6000 96 GB (BF16 weights are 49 GB).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)decosa-api judgment module (decosa_api/verticals/judgment)
0 GBProof: partial
Judge: one temperature-0 call per question, plus seeded samplesQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Fast judge with option-token probabilities (alternate)Gemma-4-26B-A4B-itgoogle/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
26B (4B active) · 49 GBNo proof yetSelf-host only
Second judge for disagreementDeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab)
552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only
Third judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • Yes/no eval: 300 validation questions to fit calibration, 1,000 others held out.

  • One-of-four eval: 300 validation questions to fit calibration, 1,000 test questions held out.

  • Eval mode, rubric: expert 1-5 ratings of 16 summaries per article on four criteria. 10 articles for dev, 25 held out.

  • Eval mode, pairwise: expert votes between two chat answers, compared with our judge and with the GPT-4 judge on the same rows. 600 votes held out.

  • scripts/judgment_eval.py and docs/evals/typed-judgment.mdApache-2.0

    Rebuilds the splits, runs the model, fits calibration on dev only, and writes the metrics, the repeat-agreement checks and the calibration file.

  • POST /judgment/verifyApache-2.0

    Checks a record's signature against this server's key and, if you send it, the context hash. The console also checks the signature in your browser with WebCrypto.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /judgment/info, /judgment/samples; POST /judgment/judge (SSE or JSON), /judgment/eval, /judgment/batch, /judgment/v1/systemone, /judgment/verify. Keeps no context.

  • vLLM (judge):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly with logprobs (self-host).

Hardware

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code tool, not run here for this tool.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran on this card, shared with other services.

Latency per lane

  • one question, one call, direct route, card shared3.5 s

    Measuredmeasured on our server 2026-09-25: median of the BoolQ test calls (about 270 prompt tokens by vLLM's count), 16 in flight, other evals on the same card

  • support-ticket sample (4 questions + 4 labels, samples k=4: 35 calls), hosted gateway route35.0 s

    Measuredmeasured on our server 2026-09-25 while the shared card had 30-70 requests queued; the calls of one request run 8 at a time, so a quiet card is several times faster

Notes

  • Logprobs are clearly the best probability: on held-out BoolQ and MMLU they rank right against wrong answers far better (AUROC 0.857 and 0.863) than 4 samples (0.651 and 0.768) or the model's stated confidence (0.703 and 0.602), which is almost always 95 to 100. Samples are well calibrated on average (ECE 0.023 on BoolQ) but coarse. That gap is the case for the gateway logprobs passthrough.
  • Re-runs are not bit-identical on a shared, busy server: batching changes the arithmetic. Repeating the temperature-0 call on 1,000 BoolQ and 1,000 MMLU test questions kept 99.0% and 95.7% of answers; every change was on an answer the calibrated logprobs put under 0.85, and none of the 1,274 answers at 0.9 or more changed. Through the hosted route (answer plus 4 seeded samples, 400 questions, run twice) 98.5% and 97.5% of answers held, 93% and 83% of distributions were identical, and none of the 182 answers at 0.9 or more changed. The API marks answers under 0.75 as near_tie. Pinned weights and prompts rule out the other kind of drift: a closed model being replaced.
  • Measured through the gateway at its list price for Qwen3.8-27B ($0.30 per million prompt tokens, $1.50 per million generated): a yes/no or choice question with a ~300-token context costs about $0.12 per 1,000 judgments with one call and no reason, about $0.17 with a reason, and $0.58-0.61 per 1,000 with the default 4 samples (5 calls, each paying the prompt). Self-hosted with logprobs: one call, GPU time only. Batches of up to 50 items report the exact cost per item.
  • Eval mode on held-out data: on 25 SummEval articles (400 summaries, four criteria) the expected 1-5 score has a mean per-article Spearman correlation of 0.525 with expert ratings (logprobs; 0.491 with samples). On 600 MT-Bench expert votes the pairwise judge agreed 64.8% with ties and 81.3% without ties; the GPT-4 judge on the same rows agreed 65.2% and 83.6%. Both orders picked the same winner 92% of the time.
  • Calibration was fit on the dev splits only (a logprob temperature, a vote prior and a stated-confidence table per answer type); the test splits were run once. Scores (1-5) use a temperature and prior chosen on SummEval dev articles; labels use the yes/no settings and are not separately measured.
  • The gateway strips logprobs today. docs/proposals/gateway-logprobs.patch in decosa-api is a proposed 40-line change to the gateway that forwards them on non-streaming calls; the existing gateway tests pass with it, but it is not applied.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

typed-judgment/assemble-prompt.md126 lines
# Assemble the Decosa typed-judgment API on this machine

You are setting up a typed-judgment service: it answers narrow questions about a context (yes/no, one of N options, a
score on described levels, or a set of labels) with a typed answer, a probability, a short reason and a record signed by
this box's own key. It also grades an output against a rubric, or compares two outputs. Work step by step, show me each
command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/typed-judgment.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py typed-judgment` (the api image carries the same bundle under /app/rehearsal/typed-judgment/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py typed-judgment --bundle typed-judgment.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer asks for a refund: yes", "the refund answer carries a probability of at least 0.8", "the returns team handles it first"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), weights `nvidia/Qwen3.8-27B-NVFP4` at revision
  `482ca0f3832238542f8f5295dde86b5f22711d80`. The engine is decosa-api (AGPL-3.0-or-later) and needs no GPU of its own.
- Pin everything: the weights revision, the vLLM image, and the decosa-api release. Reproducible verdicts are the point;
  an unpinned `latest` breaks them.
- Contexts stay on this machine. Bind every port to 127.0.0.1. The service keeps no context or question text: logs carry
  counts only. Keep it that way; do not add request logging.
- Be honest about what it does: the probabilities are estimates calibrated on public benchmarks, not guarantees.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
   6000 96 GB or an RTX 5090 32 GB both work). Driver 570 or newer. Blackwell cards run NVFP4; on older cards use
   `Qwen/Qwen3.8-27B-FP8` and expect slightly different probabilities (re-check calibration, step 4.6).
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that contains
  `decosa_api/verticals/judgment/` (`main` until one does: v0.1.0 predates it), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` (vLLM 0.29.0).
- Download the weights once at the pinned revision:
  `huggingface-cli download nvidia/Qwen3.8-27B-NVFP4 --revision 482ca0f3832238542f8f5295dde86b5f22711d80`.

## 3. docker-compose.yml
Write this in `~/decosa/judgment/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--revision", "482ca0f3832238542f8f5295dde86b5f22711d80",
              "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768", "--enable-prefix-caching",
              "--max-logprobs", "20"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_JUDGMENT_MAX_CONCURRENT: "4"
      DECOSA_JUDGMENT_WORKERS: "8"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/judgment/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.

On the direct route the service reads the answer token's log-probabilities, so `"method": "auto"` uses `logprobs`: one
call per question and the best-calibrated method in our eval. Prefix caching matters: all questions of a request share the
system prompt and the context, so the server reuses that part.

On the first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private.
Never print it. Every model call gets a receipt signed with that key (status `attested`): an attestation by me, the
operator, not a proof of computation.

## 4. Smoke test
1. `curl -s localhost:8445/judgment/info | jq '{methods, model}'`: `methods.auto` must be `logprobs` and
   `model.weights_root` must be `2fafb36533890fbcb49931741eee2065c3e3662c50d4bd08da7d3bf3a1091894`.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"typed-judgment"}' | jq -r .token)`.
3. `curl -s localhost:8445/judgment/samples | jq '.[0].request | del(.method, .samples)' > req.json` (the support ticket),
   then `curl -s -XPOST localhost:8445/judgment/judge -H "authorization: Bearer $T" -H 'content-type: application/json' -d @req.json > out.json`.
   Expect `refund: yes`, `team: returns`, the `damaged` and `double-charge` labels, and `method: logprobs` on each answer.
4. Run the same request again and compare `jq .record.answers_sha256 out.json` between runs. Report whether they match.
   On a busy server the verdicts should match; the probabilities can move by a point or two (batching; on 25 Sep 2026 we
   saw 0.986 then 0.997 for the same answer), which changes `answers_sha256` but not `verdicts_sha256`.
5. `jq '{record}' out.json | curl -s -XPOST localhost:8445/judgment/verify -H 'content-type: application/json' -d @-` must
   show `valid_signature: true` and `signed_by_this_server: true`. Change one answer in the record and verify again: it
   must fail.
6. Optional, to check calibration on this box: the eval script `scripts/judgment_eval.py` (in the decosa-api source)
   rebuilds the BoolQ and MMLU splits and prints accuracy, ECE and Brier; compare with the Stack tab.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /judgment/judge`,
`/judgment/eval` and `/judgment/batch` from your code. Harnesses written for the local decision server can switch by
pointing at `POST /judgment/v1/systemone`. Contract: `API_CONTRACT.md`, section "Typed judgments".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds confidential
contexts. If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A calibrated LLM judge API: send a context and up to 16 narrow questions, and each comes back as a typed answer (yes or no, one of N, a score or labels) with a probability you can set a threshold on, a one-line reason and a receipt, on pinned open weights so the same request can be re-run later.
Who it's for
Teams in software and ai ops.
Where it runs
Hosted or self-host
Key numbers

On 1,000 held-out BoolQ questions it scored 90.7% with expected calibration error 0.023 (logprobs, self-host route); the hosted route, which samples, scored 90.5% with ECE 0.032 on 200. Calibration was measured only on public benchmarks and varies by domain, so check it on your own labelled data.

  • 90.7% / 0.023 BoolQ accuracy / ECE, logprobs (direct route) (held out, n = 1000)
  • 83.4% / 0.026 MMLU accuracy / ECE, logprobs (direct route) (held out, n = 1000)
  • 90.5% / 84.0% Hosted route, samples k=4: BoolQ / MMLU accuracy (held out, n = 200)
  • 4.1 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Hosted or self-host
Checks
Receipt per model call; signed record with the weights, prompt and seed
Output
Structured data · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What does the Decosa typed-judgment API return?

Send the Decosa typed-judgment API a context and up to 16 narrow questions. Each comes back as a typed answer (yes or no, an option key, a level or a set of labels) with a probability, the full distribution, a one-line reason and the receipt of every model call; the request ends with a signed record of the weights, prompt, seed and answers. The answers are model outputs with estimated probabilities, not decisions: keep the decision rule in your own code.

What is LLM calibration, and is this LLM judge calibrated?

A calibrated judge's probabilities match how often it is right: of the answers it gives 0.8, about 80% should be correct, and expected calibration error (ECE) measures the gap. On 1,000 held-out BoolQ questions this API scored 90.7% with ECE 0.023 using logprobs, and 83.4% with ECE 0.026 on MMLU. The hosted route, which uses four seeded samples, scored 90.5% and 84.0% on 200 questions each (ECE 0.032 and 0.028). Logprobs need self-hosting today, stated confidence is almost always 95-100, and calibration varies by domain, so check it on your own labelled data.

Will the same request give the same answer next quarter?

The Decosa typed-judgment API pins open weights and records the weights, prompt and seed, so a request can be re-run later against the same model, which a closed judge that is updated or retired cannot offer. Repeated temperature-0 calls gave the same answer 99.0% of the time on BoolQ and 95.7% on MMLU. In our re-runs no answer with a probability of 0.9 or more changed; a busy server is not bit-reproducible, and near-ties are flagged.

How good is an LLM judge compared with human ratings?

Mixed, and it depends on the task. In eval mode, which grades an output against a rubric or compares two outputs in both orders, its mean Spearman correlation with the expert mean on SummEval was 0.525 over 25 held-out articles. On MT-Bench it agreed with expert pairwise votes 64.8% of the time with ties and 81.3% without; GPT-4 as pair judge scored 65.2% on the same rows. Pairwise probabilities are poorly calibrated (ECE 0.17).

Can I use typed judgments in hiring or credit decisions?

Where a Decosa typed judgment feeds a decision about a person, rules such as NYC Local Law 144, Illinois HB 3773 (in force 1 Jan 2026), Colorado's AI Act (effective 1 Jan 2027) and the EU AI Act's high-risk duties (Annex III uses from 2 Dec 2027) can apply to you as the deployer. The signed records and receipts help with record-keeping but are not a compliance programme. Keep human review in your own process. Not legal advice.

What does a typed-judgment request cost?

A four-question support-ticket request to the Decosa typed-judgment API took 35 calls and about 13,000 tokens with four samples, $0.004 at the gateway list price, or 7 calls and about 2,800 tokens with stated confidence, $0.001. It took about 4 s on a quiet GPU and 25-80 s while the shared GPU was busy (25 Sep 2026). The hosted demo keeps no context; for personal or confidential data, self-host.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Typed-judgment API

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.