Check if an open model can take over your prompt
Replay prompts you've already logged on an open model. In minutes you see whether it gives the same answers, which examples differ and why, and what 1,000 requests would cost: go, no-go or not enough data yet.
Built on: Typed judgment, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the open-model migration check API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the candidate and judge; scoring runs on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- migration-check
Use the hosted API
# Decosa open-model migration check: use the hosted API
You are wiring Decosa's migration check into this project. It answers "can an open model take over this prompt?": it
runs a production prompt template on an open model (Qwen3.8-27B) over examples we already logged from our closed model,
scores each output against the one we logged, and returns agreement with a 95% interval, the failure clusters, latency,
cost per 1,000 requests and a go / no-go, sealed in a signed record. It never calls a closed API and never needs our
closed-model key: we send outputs we already have. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- Agreement with our current model is not correctness. Where we have human labels, send them too.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "migration-check"}` returns `{"token", "expires_at", "budget"}`.
a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`); up to 40 examples per run (200 with a key).
3. The server estimates the generated tokens a run needs and answers 402 if the budget is short. One run at a time per token.
## The run
`POST /migration/runs` (token). Body:
```json
{
"title": "Ticket triage",
"task": {"type": "json", "schema": {"type": "object", "required": ["category"], "properties": {"category": {"enum": ["billing", "other"]}}}},
"prompt": {"system": "You triage tickets...", "template": "Ticket:\n{{input}}"},
"examples": [{"id": "t1", "input": "I was charged twice", "output": "{\"category\": \"billing\"}", "label": null}],
"incumbent": {"model": "gpt-4o-mini"},
"thresholds": {"pass_rate": 0.9, "schema_valid": 0.98},
"max_output_tokens": 600
}
```
- `task.type`: `exact` (same answer after normalising whitespace, case, quotes, code fences, a trailing full stop),
`label` (`labels` list, `extract`: `whole`, `first`, `last` or `double_bracket` for `[[X]]`), `json` (`schema`: a JSON
Schema subset, `fields`: dotted paths to compare; default the schema's top-level properties), `freetext` (`rubric`:
what matters, sent to the judge).
- `examples`: 10 to 200. `input` is a string (fills `{{input}}`) or an object of named fields for named placeholders.
`output` is what our current model returned. Optional per example: `label` (the right answer), `latency_ms`,
`prompt_tokens`, `completion_tokens` from our logs (then cost and latency use them instead of estimates).
- `incumbent.model`: matched against the dated price table (`GET /migration/info`); or send
`incumbent.price: {"input_per_m": 0.15, "output_per_m": 0.60}` in dollars per million tokens.
- Verdict: `go` when the whole 95% interval of the pass rate clears `pass_rate`; `no-go` when it sits wholly below;
`inconclusive` otherwise, or when more than 10% of examples could not be scored. Failed calls count as fails.
- JSON by default: `{report_id, headline, recommended, verdicts, candidates: {<id>: summary}, identity, examples, receipts, report, budget}`.
Each summary: `pass {k, n, rate, ci95}`, `clusters [{cluster, count, share, examples}]`, `schema_valid`, `fields`,
`labels {current, candidate}`, `confusion`, `judge_outcomes`, `latency_ms`, `cost.per_1k`, `verdict`, `reasons`,
`stability.bootstrap_same_verdict`. Numbers are strings; there are no floats in a signed record.
- With `Accept: text/event-stream` (or `"stream": true`) it streams `ready`, then `receipt` and `example` events as each
example is scored (not in input order), `identity`, `summary`, `record`, `report`, `budget`, `done`.
Other endpoints (no token): `GET /migration/info` (tasks, limits, the judge prompt and its sha256, the price table),
`GET /migration/samples`, `POST /migration/verify` `{"record": {...}}` → `{ok, record: {...}, recompute: {ok, diffs}}`,
`POST /record/verify`, `GET /attest/signing-key`, `GET /receipts/{id}`.
## Example: check a label prompt from our logs (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
rows = [json.loads(l) for l in open("logs/intent-router.jsonl")][:200]
body = {"task": {"type": "label", "labels": ["billing", "tech", "sales", "other"], "extract": "whole"},
"prompt": {"system": open("prompts/router.txt").read(), "template": "{{input}}"},
"examples": [{"id": r["id"], "input": r["message"], "output": r["model_output"], "label": r.get("agent_label")} for r in rows],
"incumbent": {"model": "gpt-4.1-mini"}, "thresholds": {"pass_rate": 0.95}, "max_output_tokens": 20}
r = httpx.post(f"{API}/migration/runs", json=body, headers=H, timeout=900)
r.raise_for_status()
rep = r.json()
print(rep["headline"])
for c in rep["candidates"].values():
print(c["verdict"], c["pass"], c["cost"]["per_1k"])
for cl in c["clusters"][:5]:
print(" ", cl["count"], cl["cluster"], cl["examples"])
json.dump(rep["report"]["record"], open("migration-record.json", "w"))
```
## Keep and verify the record
`report.record` is a `decosa.record.v1` hash chain signed with the server's Ed25519 key: the prompt, every rendered input,
both outputs, every score and judgment (each model call with its gateway receipt) and the numbers. Strip the `text` field
from every entry to share it without our data; it still verifies. `POST https://api.decosa.ai/migration/verify` checks the chain,
the signature and the receipts, and recomputes every stated number from the per-example entries.
## Honest limits
- The free-text judge is a language model. Its agreement with human experts on MT-Bench is on the Stack tab; it is
close to GPT-4's own, not perfect.
- Latency is measured on a shared GPU through the gateway; our own deployment will differ.
- Closed-model cost uses list prices checked on the date shown and estimated token counts unless we send ours.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa open-model migration check: run it yourself (containers)
You are setting up the Decosa migration check on this machine, so our logged production prompts and outputs never leave
it. It runs our prompt on an open model, scores it against the outputs we logged from our closed model, and signs a
go / no-go record. Nothing is sent to Decosa's hosted API, and no closed API is called.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/migration-check.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py migration-check` (the api image carries the same bundle under /app/rehearsal/migration-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py migration-check --bundle migration-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 10 examples were scored", "at least 9 of 10 outputs parse against the schema", "no model call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
127.0.0.1. To compare more open models, run them as extra OpenAI-compatible services and list them in
`DECOSA_MIGRATION_CANDIDATES` (a JSON list of `{"id", "label", "base_url", "model", "license"}`); requests then pick
them with `"candidates": [...]`.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/migration/info` lists the task types, the candidates and the judge;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify our records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"migration-check"}`, fetch `GET /migration/samples`, and
send the `tickets-json` sample's `request` to `POST /migration/runs`. Expect 20 examples scored, `schema_valid` 20 of
20, a few `field priority differs` clusters and a verdict. Then `POST /migration/verify` with `report.record`: `ok`,
`record.ok` and `recompute.ok` must all be true.
6. Report back: the public key and key id, the smoke-test verdict and clusters, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
production logs. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/migration-check-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (57.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"migration-check"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py migration-check
Download the mock-data bundle (3 KB, 10 checks)expected.json
Ten fictional support tickets for a made-up homeware shop, each with the JSON a hand-written reference gave (category, priority, refund_requested, order_id). The open model answers the same prompt; its outputs must parse against the schema and mostly agree field by field, and the sealed record must verify, recompute every stated number and fail once one example's pass mark is changed.
What the rehearsal checks
- all 10 examples were scored
- at least 9 of 10 outputs parse against the schema
- no model call failed
- at least 7 of 10 outputs agree with the reference on every field
- the verdict is one of go, no-go or inconclusive
- the sealed record verifies
- every stated number recomputes from the per-example entries
- the record was issued by this server
- a record with one example's pass mark changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic: 20 fictional support tickets for a made-up shop (the first 10 here). The reference outputs were written for this demo, not produced by a closed API. Written for the Decosa demo, 25 Sep 2026. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa open-model migration check: run it yourself (containers)
You are setting up the Decosa migration check on this machine, so our logged production prompts and outputs never leave
it. It runs our prompt on an open model, scores it against the outputs we logged from our closed model, and signs a
go / no-go record. Nothing is sent to Decosa's hosted API, and no closed API is called.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/migration-check.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py migration-check` (the api image carries the same bundle under /app/rehearsal/migration-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py migration-check --bundle migration-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 10 examples were scored", "at least 9 of 10 outputs parse against the schema", "no model call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
127.0.0.1. To compare more open models, run them as extra OpenAI-compatible services and list them in
`DECOSA_MIGRATION_CANDIDATES` (a JSON list of `{"id", "label", "base_url", "model", "license"}`); requests then pick
them with `"candidates": [...]`.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/migration/info` lists the task types, the candidates and the judge;
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify our records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"migration-check"}`, fetch `GET /migration/samples`, and
send the `tickets-json` sample's `request` to `POST /migration/runs`. Expect 20 examples scored, `schema_valid` 20 of
20, a few `field priority differs` clusters and a verdict. Then `POST /migration/verify` with `report.record`: `ok`,
`record.ok` and `recompute.ok` must all be true.
6. Report back: the public key and key id, the smoke-test verdict and clusters, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
production logs. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/migration-check-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsOpen-model migration check on GeForce RTX 5090: use the Standard · Qwen3.8-27B as candidate and judge (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the grounding check, not run here for this vertical.
Standard · Qwen3.8-27B as candidate and judge (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Checker: decosa-api migration module (decosa_api/verticals/migration). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Candidate and free-text judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Open-model migration check, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Open-model migration check on my hardware Fetch https://decosa.ai/prompts/migration-check-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=migration-check) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · Qwen3.8-27B as candidate and judge (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Checker: decosa-api migration module (decosa_api/verticals/migration), CPU - Candidate and free-text judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/migration-check-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Open-model migration check: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Open-model migration check on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/migration-check.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py migration-check` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 10 examples were scored", "at least 9 of 10 outputs parse against the schema", "no model call failed"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Checker: rendering, scoring, statistics, cost, signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Candidate and free-text judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py migration-check`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 3.6 s · ~$0.003 per run · 30 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, compose up, sample against local model servers
Measured cost to run: about $0.015 per 100 examples (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. 20 synthetic tickets: schema valid 20 of 20, a verdict with reasons; the record verifies and recomputes, and a flipped score fails at that entry with the numbers that no longer follow named. Key minting with the admin secret works.
Known limits (3)
- The hosted numbers are for the 20-ticket JSON sample. The 20-question MT-Bench free-text sample (70 calls with the judge) took 19 s on a quiet GPU and 3-7 minutes while the shared GPU was busy (25 Sep 2026).
- Agreement with your current model is not correctness; send human labels where you have them.
- Latency in the report is measured on a shared GPU through the gateway; your own deployment will differ.
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the open-model migration check API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the candidate and judge; scoring runs on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Can an open model take over this prompt? A shadow run against the outputs you already logged, with a go / no-go in a signed record.
Send a production prompt template and 10 to 200 examples with the outputs your closed model already returned. The same inputs run on an open model, one receipted call each. Structured tasks are scored in code (exact answer, label, JSON schema and field by field); free text goes to a pinned open judge that compares both answers in both orders. You get agreement per example with a 95% interval, the failure clusters, accuracy against your human labels where you have them, latency, cost per 1,000 requests against the dated list price, and a go, no-go or inconclusive verdict. The whole run is sealed into a signed record whose numbers anyone can recompute. We never call the closed API and never need its key.
- Deployment
- Hosted or self-host
- Regulatory
- A migration check compares outputs; it does not certify a model. Agreement with your current model is not correctness, the free-text judge is a language model and can be wrong, and the interval covers only inputs like the ones you sent. Closed-API terms can restrict what you do with outputs: Anthropic's Commercial Terms (section D.4, effective 17 Jun 2025) forbid using the services 'to train competing AI models'. This check trains nothing and only compares, but read your own provider's terms. Production logs often hold personal or customer data, which data-protection law (GDPR, CCPA and others) and your customer contracts govern: the hosted demo keeps nothing, and real logs belong on a self-hosted box. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B, Gemma-4-31B-it), MIT (DeepSeek-V4-Flash-0731). Prices and terms checked 25 Sep 2026.
Text description
A prompt template, the examples with the outputs your closed model already returned, and a pass bar go to the checker; the closed API is never called. The checker renders each prompt and sends it to the open candidate, Qwen3.8-27B, one receipted call per example. Exact, label and JSON tasks are scored in code; free text goes to the same pinned model as a judge, which compares both answers in both orders. The endpoint auditor's golden prompts check that the hosted candidate is the claimed model. The checker computes agreement with a 95% interval, failure clusters, latency and cost per 1,000 requests, decides go, no-go or inconclusive, and seals it all into a signed record. Outputs: per-example results side by side, the verdict, and the record. On the hosted route each model call gets a gateway-signed receipt. In self-host mode the checker and the models run on your machine.
At a glance
- Data retention
- Nothing stored: prompts, inputs and outputs live in memory for the request and come back to you in the sealed record. Strip each entry's text to share the record without your data; it still verifies.
- What leaves the box
- Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box. It never calls a closed API and never needs your closed-model key.
- Input formats
- JSON: a prompt template and 10 to 200 logged examples (up to 40 with a demo session), each with the current model's output and optionally a human label and logged tokens and latency. Tasks: exact, label, JSON (schema) or free text (judge).
- Typical run
- JSON tickets: a fraction of a cent at the gateway list price. Free-text answers with the judge in both orders: more calls, a few cents. Each run shows its own measured cost.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
structured prompts on a 24-32 GB card (self-host)
Exact, label and JSON prompts need no judge: compare a small open model such as Gemma-4-31B against your logs. No free-text judging on this tier.
- Models
- decosa-api migration module (decosa_api/verticals/migration)
- Gemma-4-31B-it
- Hardware
- 1x RTX 4090 24 GB or RTX 5090 32 GB (estimate)
- Quality evidence
- Structured scoring (exact, label, JSON schema, fields)deterministic code, covered by unit tests; no model to measuretests/test_migration.py
- Gemma-4-31B as a candidatenot measured yetnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyNot a hosted model; calls get receipts signed by your own box (attested).
- In the hosted demo
Standard
Qwen3.8-27B as candidate and judge (hosted demo)
One pinned open model runs your prompt and judges free text in both orders. This is what the hosted API runs.
- Models
- decosa-api migration module (decosa_api/verticals/migration)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
- Quality evidence
- MT-Bench test, agreement with human experts on 'is the candidate worse?' (both orders)79.3% (95% CI 74.4-83.5), κ 0.588; runs 2 and 3: 77.3%, 78.3%docs/evals/migration-check.md, 300 held-out pairs
- Same items, GPT-4 as judge (published MT-Bench verdicts)78.0% (73.0-82.3), κ 0.562docs/evals/migration-check.md
- Agreement without ties (judge vs experts)89.5% (GPT-4: 88.0%)docs/evals/migration-check.md
- Report verdict unchanged across 3 runs / matches the experts' verdict14 of 15 pairings / 14 of 15 (GPT-4: 13 of 15)docs/evals/migration-check.md
- Latency
- measured on a shared GPU: seconds per output under load; under a minute for a JSON run and a couple of minutes for a free-text run.
- Verification
- Proof: strongEvery candidate output and judge call has a gateway-signed receipt embedded in the signed record.
Best
DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host)
A 284B-parameter MoE for prompts the 27B cannot carry. Uses both cards, so the judge needs another box or runs between passes.
- Models
- decosa-api migration module (decosa_api/verticals/migration)
- DeepSeek-V4-Flash-0731
- Hardware
- 2x RTX PRO 6000 96 GB with community vLLM patches
- Quality evidence
- As a migration candidatenot measured yetnot measured yet
- Latency
- not measured for this vertical; the owner measured 109-151 t/s single-stream on this box
- Verification
- No proof yetSelf-host onlyNot a hosted model; attested receipts only.
- Needs more compute
Wanted: the best setup
the largest open candidates
Check a migration against DeepSeek-V4.1-Flash and GLM-5.3-Flash, the most-used open models, with Qwen3.8-27B as the judge. Not served yet.
- Models
- decosa-api migration module (decosa_api/verticals/migration)
- Qwen3.8-27B (NVFP4)
- DeepSeek-V4.1-Flash
- GLM-5.3-Flash
- Hardware
- Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.
- Quality evidence
- agreement with expert labels, same protocolnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyNot hosted yet, so no receipts today.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)decosa-api migration module (decosa_api/verticals/migration) 0 GBProof: partial | LiteStandardBestWanted | 0 GB | Proof: partial | |
| ||||
Candidate and free-text judgeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | StandardWanted | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Small candidate for structured tasks (self-host)Gemma-4-31B-itgoogle/gemma-4-31B-it on Hugging Face (opens in a new tab) 31B · 20 GBNo proof yetSelf-host only | Lite | 31B · 20 GB | No proof yetSelf-host only | |
| ||||
Large candidate (self-host, two GPUs)DeepSeek-V4-Flash-0731deepseek-ai/DeepSeek-V4-Flash-0731 on Hugging Face (opens in a new tab) 284B (13B active) · 176 GBNo proof yetSelf-host only | Best | 284B (13B active) · 176 GB | No proof yetSelf-host only | |
| ||||
Candidate: the largest open flash modelDeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab) 552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only | Wanted | 552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate) | No proof yetSelf-host only | |
| ||||
Candidate: the most-used open modelGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) 321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only | Wanted | 321B (18B active) · about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
The eval and two demo sets: expert pairwise votes on answers from GPT-4, GPT-3.5-turbo, Claude-v1 and three open models, plus GPT-4's own verdicts as a baseline judge.
The 'production prompt' in the pairwise-judge demo (GPT-4 as the current model).
- scripts/migration_eval.py and docs/evals/migration-check.mdApache-2.0
Builds the dev and test splits by question, runs the judge in both orders, repeats it, and writes agreement with the human experts against GPT-4's.
- POST /migration/verifyApache-2.0
Checks a record's chain, signature and receipts, and recomputes every stated number from its per-example entries. The console also checks the signature in your browser.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /migration/info, /migration/samples; POST /migration/runs (SSE or JSON), /migration/verify. Stores nothing.
- vLLM (candidate and judge):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX 5090 32 GB Fits
Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the grounding check, not run here for this vertical.
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo and the eval ran on this card, shared with other services.
- 2x RTX PRO 6000 96 GB Fits
Needed for the DeepSeek-V4-Flash candidate (both cards); the owner has served it with community patches. Not run for this vertical.
Latency per lane
- one candidate output, JSON triage (about 35 tokens), 6 in flight, hosted gateway route8.1 s
Measuredmeasured on our server 2026-09-25: p50 8.1 s, p95 11.4 s over 20 calls while the GPU was shared with other evaluation jobs (queueing dominates; a single quiet call of a few tokens took 0.2 s the same day)
- one candidate output, MT-Bench answer (about 390 tokens), hosted gateway route8.8 s
Measuredmeasured on our server 2026-09-25: p50 8.8 s, p95 15.3 s over 20 calls, shared GPU
- whole demo run, 20 free-text examples with both judge orders (60 model calls)115.0 s
Measuredmeasured on our server 2026-09-25, shared GPU; 20 JSON examples took 33 s
- one judge call, 8 in flight, eval14.6 s
Measuredmeasured on our server 2026-09-25: mean 13.5-20.8 s per call across three runs of 615 calls, GPU saturated by other work
Notes
- On 300 held-out MT-Bench pairs the judge agreed with the human experts on whether the candidate was worse 79.3% of the time (κ 0.588); GPT-4, MT-Bench's own judge, agreed 78.0% (κ 0.562) on the same items. The intervals overlap: on par, not better.
- Running the judge in both orders is the default: it costs a second call and raised agreement from 76.3% to 79.3% and precision on 'worse' from 0.69 to 0.75.
- Repeatability: across three runs 84.3% of items got the identical verdict; at report level (15 model pairings, 0.8 bar) 14 of 15 got the same go, no-go or inconclusive every time, and 14 of 15 matched the verdict the expert labels give. When it misses, it is usually stricter than the people.
- The prompt was written once and checked on a dev split of 24 questions; the 56-question test split was never used for changes.
- Agreement with the current model is not correctness. The pairwise-judge demo shows it: Qwen agrees with GPT-4 on 18 of 24 verdicts, yet against the human labels GPT-4 is right on 16 and Qwen on 15. Send human labels where you have them.
- Cost uses dated list prices: at open-market prices Qwen3.8-27B is far cheaper than GPT-4-class or Claude Sonnet-class models, but not cheaper than the smallest closed tiers (gpt-5.6-luna, gpt-4o-mini, gemini-2.5-flash-lite). The report shows the number either way.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa open-model migration check on this machine
You are setting up a migration checker: it takes a production prompt template and examples we logged from our closed
model (input and output, optionally a human label), runs the same inputs on open models here, scores each output against
the logged one, and returns agreement with a 95% interval, failure clusters, latency, cost per 1,000 requests and a
go / no-go, sealed into a record signed by this box's own key. Work step by step, show me each command before you run
anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/migration-check.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py migration-check` (the api image carries the same bundle under /app/rehearsal/migration-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py migration-check --bundle migration-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 10 examples were scored", "at least 9 of 10 outputs parse against the schema", "no model call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) as the default candidate and as the free-text judge. Any extra candidate you add
must have a licence that allows our use; write it down. The checker is decosa-api (AGPL-3.0-or-later) and needs no GPU.
- Our prompts and logged outputs stay on this machine. Bind every port to 127.0.0.1. The service stores nothing: runs
live in memory and the signed record goes back to the caller; logs carry counts only. Do not add request logging.
- Never call a closed API from this setup and never ask me for a closed-model key. The closed model's outputs are the
ones we already logged.
- Be honest about what it measures: agreement with our current model is not correctness, and the judge is a language
model that can be wrong.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
6000 96 GB or an RTX 5090 32 GB both work). Driver 570 or newer. Blackwell cards run NVFP4; on older cards use the FP8
weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free for one candidate; more for each extra candidate.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that contains
`decosa_api/verticals/migration/` (`main` until one does: v0.1.0 predates it), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8` on a card
without NVFP4). Engine and quant change output quality on this model, so keep the pinned image and weights.
## 3. docker-compose.yml
Write this in `~/decosa/migration/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching", "--seed", "0"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_ADMIN_SECRET: ${DECOSA_ADMIN_SECRET} # for minting keys on this box; put it in ~/decosa/migration/.env (0600)
DECOSA_MIGRATION_MAX_CONCURRENT: "2"
DECOSA_MIGRATION_WORKERS: "6"
# extra candidates (optional), each an OpenAI-compatible server you run:
# DECOSA_MIGRATION_CANDIDATES: '[{"id": "gemma-4-31b", "label": "Gemma-4-31B", "base_url": "http://gemma:8000/v1", "model": "gemma-4-31b", "license": "Apache-2.0"}]'
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/migration/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.
Demo sessions are capped at 40 examples and 20,000 generated tokens. For real runs (up to 200 examples) mint a key on
this box: `curl -s -XPOST localhost:8445/v1/keys -H "authorization: Bearer $DECOSA_ADMIN_SECRET" -H 'content-type: application/json' -d '{"label": "migration", "verticals": ["migration-check"], "daily_llm_tokens": 1000000}'`
and keep the returned `dk_…` key in an environment variable. A free-text run of 200 examples needs roughly 200,000
generated tokens (the candidate's outputs plus two judge calls each); a JSON or label run far fewer.
On the first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private.
Never print it. Every model call on the direct route gets a receipt signed with that key (status `attested`): an
attestation by me, the operator, not a proof of computation. The signed record uses the same key.
## 4. Smoke test
1. `curl -s localhost:8445/migration/info | jq '{candidates, judge: .judge.prompt_sha256, prices: .prices.version}'`.
2. Get a token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"migration-check"}' | jq -r .token)`.
3. `curl -s localhost:8445/migration/samples | jq '.[] | select(.id=="tickets-json") | .request' > tickets.json`, then
`curl -s -XPOST localhost:8445/migration/runs -H "authorization: Bearer $T" -H 'content-type: application/json' -d @tickets.json > run.json`.
Expect 20 examples, `schema_valid` 20 of 20 and a verdict with its reasons. Priority is the field most likely to
differ: those tickets are judgement calls.
4. Stream the same body with `-H 'accept: text/event-stream' -N`: `ready`, then `receipt` and `example` events,
`identity`, `summary`, `record`, `report`, `budget` and `done`.
5. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/migration/verify -H 'content-type: application/json' -d @-`
must show `ok: true`, `record.ok: true` and `recompute.ok: true`. Then flip one `"pass"` in a `score` entry and verify
again: it must fail at that entry, and the recompute must name the numbers that no longer follow.
6. Run our own prompt: export 20-200 logged rows as `{"id", "input", "output"}` (plus `label` where a person checked the
answer, and `latency_ms`, `prompt_tokens`, `completion_tokens` if the logs have them) and send them the same way.
Tell me the verdict, the top three clusters and the cost line.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local` to use the console against this box, or
call `POST /migration/runs` from a script in CI whenever the prompt changes, and keep each signed record with the
prompt's version. Contract: `API_CONTRACT.md`, section "Open-model migration check".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds production logs.
If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Can an open model take over this prompt? A shadow run against the outputs you already logged, with a go / no-go in a signed record.
- Who it's for
- Teams in software and ai ops.
- Where it runs
- Hosted or self-host
- Key numbers
- 79.3% (74.4-83.5) Not-worse agreement with human experts, both orders (run 1) (test split, n = 300)
- 0.588 Cohen's kappa vs experts (run 1) (test split, n = 300)
- 0.755 / 0.839 Precision / recall on "worse" (run 1) (test split, n = 300)
- 3.6 s Median end-to-end run, hosted (QA sweep 2026-09-25)
- Models
- Qwen3.8-27B (candidate and judge)
- Where
- Hosted or self-host
- Checks
- Receipt per output and judgment; signed record that recomputes
- Industry
- Software and AI ops
- Input
- Text and documents
- Output
- Signed record or verdict · Structured data
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Typed judgment · Signed record
Questions people ask
How does the open-model migration check decide whether a model can take over a prompt?
The open-model migration check takes a production prompt template and 10 to 200 examples with the outputs your closed model already returned, and reruns the same inputs on an open model. Structured tasks are scored in code; free text goes to a pinned open judge that compares both answers in both orders. You get agreement per example with a 95% interval, failure clusters, cost per 1,000 requests and a go, no-go or inconclusive verdict.
Does the migration check need my OpenAI or Anthropic API key?
No. The open-model migration check never calls the closed API and never needs its key: you paste the outputs your closed model already logged. Hosted, each model call goes through our gateway to the GPU serving Qwen3.8-27B and only receipts (hashes, token counts, no text) are kept. Self-hosted on the direct route, nothing leaves the box. Read your own provider's terms on using outputs.
How accurate is the free-text judge in the migration check?
On 300 held-out MT-Bench items, the open-model migration check's judge agreed with human experts on 'not worse' 79.3% of the time (74.4-83.5), against 78.0% for GPT-4 on the same items: on par, not better. Cohen's kappa was 0.588. Caveats: MT-Bench answers are 2023-era and general-purpose, first-turn only, and self-preference when the judge scores its own answers was not measured.
Is agreement with my current model the same as being correct?
No. The open-model migration check compares outputs; it does not certify a model. Agreement with your current model is not correctness, the free-text judge is a language model and can be wrong, and the interval covers only inputs like the ones you sent. Send human labels where you have them: the report then shows accuracy against those labels as well as agreement with the incumbent.
Will an open model be cheaper than my closed model?
Not always. The value note behind the open-model migration check found Qwen3.8-27B about 5-8x cheaper than GPT-5.4 or Claude Sonnet 4.6 class models at open-weight market prices on 25 Sep 2026, but not cheaper than small closed tiers such as gpt-4o-mini or gemini-2.5-flash-lite. The report shows cost per 1,000 requests against the dated list price, so for small tiers the reason to switch is control and residency.
Can I run the migration check on real production logs?
Yes, but self-hosted. The open-model migration check keeps nothing: prompts, inputs and outputs live in memory for the request and come back in the sealed record. Production logs often hold personal or customer data governed by GDPR, CCPA and your contracts, so real logs belong on a self-hosted box. You can strip each entry's text to share the record; it still verifies.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Open-model migration check
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…