{"schema_version":"1","site":"https://decosa.ai","id":"migration-check","num":"23","name":"Open-model migration check","tool_name":"Check if an open model can take over your prompt","short":"Migration check","blurb":"Paste a production prompt and the outputs your closed model already gave. The same inputs run on an open model; you get agreement per example with an interval, the failure clusters, cost per 1,000 requests, and a go or no-go in a signed record.","status":"live","labels":{"industry":["software"],"job":["review","attest"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["record","data"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["software"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["typed-judgment","signed-record"],"models":"Qwen3.8-27B (candidate and judge)","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the candidate and judge; scoring runs on CPU","final_artifact":"A go / no-go with failure clusters, cost per 1,000 requests and a signed record anyone can re-check.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per output and judgment; signed record that recomputes","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":3603,"p95_ms":null,"runs":null,"receipts_per_run":30,"cost_per_run_usd":0.0029},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. 20 synthetic tickets: schema valid 20 of 20, a verdict with reasons; the record verifies and recomputes, and a flipped score fails at that entry with the numbers that no longer follow named. Key minting with the admin secret works."},"known_limits":["The hosted numbers are for the 20-ticket JSON sample. The 20-question MT-Bench free-text sample (70 calls with the judge) took 19 s on a quiet GPU and 3-7 minutes while the shared GPU was busy (25 Sep 2026).","Agreement with your current model is not correctness; send human labels where you have them.","Latency in the report is measured on a shared GPU through the gateway; your own deployment will differ."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Not-worse agreement with human experts, both orders (run 1)","value":"79.3% (74.4-83.5)","unit":null,"n":300,"split":"test","note":"Runs 2 and 3: 77.3% and 78.3%. GPT-4 (MT-Bench's own judge) on the same items: 78.0%. On par, not better."},{"name":"Cohen's kappa vs experts (run 1)","value":"0.588","unit":null,"n":300,"split":"test","note":"Runs 2 and 3: 0.546, 0.568; GPT-4: 0.562"},{"name":"Precision / recall on \"worse\" (run 1)","value":"0.755 / 0.839","unit":null,"n":300,"split":"test","note":"Errs on the strict side more than the lenient one."},{"name":"Verdict identical in all three runs, per item","value":"84.3%","unit":null,"n":300,"split":"test","note":"Temperature 0 on a batched server is not bit-exact; variation lands on close calls."},{"name":"Report-level verdict equal to the experts' labels","value":"14 of 15","unit":"pairings","n":15,"split":"test","note":"GPT-4: 13 of 15. 14 of 15 pairings kept the same verdict across three runs."},{"name":"Dev agreement, both orders","value":"85.0%","unit":null,"n":120,"split":"dev","note":"GPT-4 87.5% on the same dev items"}],"dataset":"lmsys/mt_bench_human_judgments (CC-BY-4.0): expert pairwise votes on first-turn MT-Bench answers from six 2023-era models, with GPT-4's own verdicts as the baseline judge. Split by question: 120 dev items, 300 test items; the prompt was written once, run once on dev, not changed, and test was run three times.","held_out":true,"caveats":["Measures only the free-text judge; the structured scorers are deterministic code covered by unit tests.","MT-Bench answers are 2023-era and general-purpose; agreement on a team's own task should be checked against a few of their own labels.","Human tie votes are noisy (about a quarter of items); not-worse folds them into \"not worse\".","First-turn answers only; no multi-turn conversations.","Self-preference when the judge model judges its own answers is not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/migration-check"},"stack":{"summary":"Send a production prompt template and 10 to 200 examples with the outputs your closed model already returned. The same inputs run on an open model, one receipted call each. Structured tasks are scored in code (exact answer, label, JSON schema and field by field); free text goes to a pinned open judge that compares both answers in both orders. You get agreement per example with a 95% interval, the failure clusters, accuracy against your human labels where you have them, latency, cost per 1,000 requests against the dated list price, and a go, no-go or inconclusive verdict. The whole run is sealed into a signed record whose numbers anyone can recompute. We never call the closed API and never need its key.","tagline":"Can an open model take over this prompt? A shadow run against the outputs you already logged, with a go / no-go in a signed record.","deployment":"hosted-or-self-host","regulatory_note":"A migration check compares outputs; it does not certify a model. Agreement with your current model is not correctness, the free-text judge is a language model and can be wrong, and the interval covers only inputs like the ones you sent. Closed-API terms can restrict what you do with outputs: Anthropic's Commercial Terms (section D.4, effective 17 Jun 2025) forbid using the services 'to train competing AI models'. This check trains nothing and only compares, but read your own provider's terms. Production logs often hold personal or customer data, which data-protection law (GDPR, CCPA and others) and your customer contracts govern: the hosted demo keeps nothing, and real logs belong on a self-hosted box. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B, Gemma-4-31B-it), MIT (DeepSeek-V4-Flash-0731). Prices and terms checked 25 Sep 2026.","components":[{"id":"checker","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","name":"decosa-api migration module (decosa_api/verticals/migration)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; JSON Schema subset validator, Wilson intervals, seeded bootstrap","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","best"],"alternative_to":null},{"id":"qwen","role":"Candidate and free-text judge","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, seed fixed, thinking off","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"gemma","role":"Small candidate for structured tasks (self-host)","name":"Gemma-4-31B-it","hf_repo":"google/gemma-4-31B-it","license":"Apache-2.0","params":"31B","quant":"Q4 or FP8 (about 18-31 GB)","vram_gb":20,"memory_gb_estimate":null,"engine":"any OpenAI-compatible server (vLLM, llama.cpp), added with DECOSA_MIGRATION_CANDIDATES","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"deepseek","role":"Large candidate (self-host, two GPUs)","name":"DeepSeek-V4-Flash-0731","hf_repo":"deepseek-ai/DeepSeek-V4-Flash-0731","license":"MIT","params":"284B","quant":"NVFP4 (about 159-176 GB)","vram_gb":176,"memory_gb_estimate":null,"engine":"vLLM TP2 with community patches on 2x RTX PRO 6000 (owner's kit), added with DECOSA_MIGRATION_CANDIDATES","receipt_coverage":"none","in_hosted_demo":false,"tiers":["best"],"alternative_to":null},{"id":"dsv41-wanted","role":"Candidate: the largest open flash model","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"glm-wanted","role":"Candidate: the most-used open model","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · structured prompts on a 24-32 GB card (self-host)","summary":"Exact, label and JSON prompts need no judge: compare a small open model such as Gemma-4-31B against your logs. No free-text judging on this tier.","components":["checker","gemma"],"hardware":"1x RTX 4090 24 GB or RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Structured scoring (exact, label, JSON schema, fields)","value":"deterministic code, covered by unit tests; no model to measure","source":"tests/test_migration.py"},{"metric":"Gemma-4-31B as a candidate","value":"not measured yet","source":"not measured yet"}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not a hosted model; calls get receipts signed by your own box (attested).","hosting":null},{"id":"standard","label":"Standard · Qwen3.8-27B as candidate and judge (hosted demo)","summary":"One pinned open model runs your prompt and judges free text in both orders. This is what the hosted API runs.","components":["checker","qwen"],"hardware":"1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"MT-Bench test, agreement with human experts on 'is the candidate worse?' (both orders)","value":"79.3% (95% CI 74.4-83.5), κ 0.588; runs 2 and 3: 77.3%, 78.3%","source":"docs/evals/migration-check.md, 300 held-out pairs"},{"metric":"Same items, GPT-4 as judge (published MT-Bench verdicts)","value":"78.0% (73.0-82.3), κ 0.562","source":"docs/evals/migration-check.md"},{"metric":"Agreement without ties (judge vs experts)","value":"89.5% (GPT-4: 88.0%)","source":"docs/evals/migration-check.md"},{"metric":"Report verdict unchanged across 3 runs / matches the experts' verdict","value":"14 of 15 pairings / 14 of 15 (GPT-4: 13 of 15)","source":"docs/evals/migration-check.md"}],"latency_note":"measured on a shared GPU: seconds per output under load; under a minute for a JSON run and a couple of minutes for a free-text run.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every candidate output and judge call has a gateway-signed receipt embedded in the signed record.","hosting":null},{"id":"best","label":"Best · DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host)","summary":"A 284B-parameter MoE for prompts the 27B cannot carry. Uses both cards, so the judge needs another box or runs between passes.","components":["checker","deepseek"],"hardware":"2x RTX PRO 6000 96 GB with community vLLM patches","quality_evidence":[{"metric":"As a migration candidate","value":"not measured yet","source":"not measured yet"}],"latency_note":"not measured for this vertical; the owner measured 109-151 t/s single-stream on this box","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not a hosted model; attested receipts only.","hosting":null},{"id":"wanted","label":"Wanted · the largest open candidates","summary":"Check a migration against DeepSeek-V4.1-Flash and GLM-5.3-Flash, the most-used open models, with Qwen3.8-27B as the judge. Not served yet.","components":["checker","qwen","dsv41-wanted","glm-wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.","quality_evidence":[{"metric":"agreement with expert labels, same protocol","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /migration/info, /migration/samples; POST /migration/runs (SSE or JSON), /migration/verify. Stores nothing."},{"name":"vLLM (candidate and judge)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"MT-Bench human judgments (lmsys/mt_bench_human_judgments)","url":"https://huggingface.co/datasets/lmsys/mt_bench_human_judgments","license":"CC-BY-4.0","purpose":"The eval and two demo sets: expert pairwise votes on answers from GPT-4, GPT-3.5-turbo, Claude-v1 and three open models, plus GPT-4's own verdicts as a baseline judge."},{"name":"FastChat pair-v2 judge prompt","url":"https://github.com/lm-sys/FastChat/blob/main/fastchat/llm_judge/data/judge_prompts.jsonl","license":"Apache-2.0","purpose":"The 'production prompt' in the pairwise-judge demo (GPT-4 as the current model)."},{"name":"scripts/migration_eval.py and docs/evals/migration-check.md","url":null,"license":"Apache-2.0","purpose":"Builds the dev and test splits by question, runs the judge in both orders, repeats it, and writes agreement with the human experts against GPT-4's."},{"name":"POST /migration/verify","url":null,"license":"Apache-2.0","purpose":"Checks a record's chain, signature and receipts, and recomputes every stated number from its per-example entries. The console also checks the signature in your browser."}],"hardware":[{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the grounding check, not run here for this vertical."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran on this card, shared with other services."},{"tier":"2x RTX PRO 6000 96 GB","fits":true,"notes":"Needed for the DeepSeek-V4-Flash candidate (both cards); the owner has served it with community patches. Not run for this vertical."}],"latency":[{"lane":"one candidate output, JSON triage (about 35 tokens), 6 in flight, hosted gateway route","typical_ms":8075,"source":"measured on our server 2026-09-25: p50 8.1 s, p95 11.4 s over 20 calls while the GPU was shared with other evaluation jobs (queueing dominates; a single quiet call of a few tokens took 0.2 s the same day)"},{"lane":"one candidate output, MT-Bench answer (about 390 tokens), hosted gateway route","typical_ms":8771,"source":"measured on our server 2026-09-25: p50 8.8 s, p95 15.3 s over 20 calls, shared GPU"},{"lane":"whole demo run, 20 free-text examples with both judge orders (60 model calls)","typical_ms":115000,"source":"measured on our server 2026-09-25, shared GPU; 20 JSON examples took 33 s"},{"lane":"one judge call, 8 in flight, eval","typical_ms":14600,"source":"measured on our server 2026-09-25: mean 13.5-20.8 s per call across three runs of 615 calls, GPU saturated by other work"}],"benchmark":null,"notes":["On 300 held-out MT-Bench pairs the judge agreed with the human experts on whether the candidate was worse 79.3% of the time (κ 0.588); GPT-4, MT-Bench's own judge, agreed 78.0% (κ 0.562) on the same items. The intervals overlap: on par, not better.","Running the judge in both orders is the default: it costs a second call and raised agreement from 76.3% to 79.3% and precision on 'worse' from 0.69 to 0.75.","Repeatability: across three runs 84.3% of items got the identical verdict; at report level (15 model pairings, 0.8 bar) 14 of 15 got the same go, no-go or inconclusive every time, and 14 of 15 matched the verdict the expert labels give. When it misses, it is usually stricter than the people.","The prompt was written once and checked on a dev split of 24 questions; the 56-question test split was never used for changes.","Agreement with the current model is not correctness. The pairwise-judge demo shows it: Qwen agrees with GPT-4 on 18 of 24 verdicts, yet against the human labels GPT-4 is right on 16 and Qwen on 15. Send human labels where you have them.","Cost uses dated list prices: at open-market prices Qwen3.8-27B is far cheaper than GPT-4-class or Claude Sonnet-class models, but not cheaper than the smallest closed tiers (gpt-5.6-luna, gpt-4o-mini, gemini-2.5-flash-lite). The report shows the number either way."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: prompts, inputs and outputs live in memory for the request and come back to you in the sealed record. Strip each entry's text to share the record without your data; it still verifies."},{"label":"What leaves the box","value":"Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box. It never calls a closed API and never needs your closed-model key."},{"label":"Input formats","value":"JSON: a prompt template and 10 to 200 logged examples (up to 40 with a demo session), each with the current model's output and optionally a human label and logged tokens and latency. Tasks: exact, label, JSON (schema) or free text (judge)."},{"label":"Typical run","value":"JSON tickets: a fraction of a cent at the gateway list price. Free-text answers with the judge in both orders: more calls, a few cents. Each run shows its own measured cost."}],"data_handling":{"page":"/data#migration-check","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: prompts, inputs and outputs live in memory for the request and come back to you in the sealed record. Strip each entry's text to share the record without your data; it still verifies.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"Extra candidate endpoints you configure","route":"selfhost","sends":"your-system","what":"Only if you list candidate models in DECOSA_MIGRATION_CANDIDATES: your examples then go to those endpoints.","default":"off","off":null}]},"console":{"href":"/tools/developer/migration-check","input":"migration","lanes":[{"id":"examples","title":"Examples","kind":"list"},{"id":"summary","title":"Agreement, clusters, cost","kind":"markdown"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"mtbench-gpt4-answers","title":"Mtbench gpt4 answers","deep_link":"/tools/developer/migration-check?sample=1&autorun=0"},{"n":2,"id":"mtbench-gpt4-judge","title":"Mtbench gpt4 judge","deep_link":"/tools/developer/migration-check?sample=2&autorun=0"},{"n":3,"id":"tickets-json","title":"Tickets json","deep_link":"/tools/developer/migration-check?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/migration-check-hosted.md","selfhost":"/prompts/migration-check-selfhost.md","assemble":"/prompts/migration-check-assemble.md","mac":"/prompts/migration-check-mac.md"},"rehearsal":{"bundle":"/samples/migration-check.zip","bundle_url":"https://decosa.ai/samples/migration-check.zip","folder":"/samples/migration-check/","expected":"/samples/migration-check/expected.json","files":["/samples/migration-check/expected.json","/samples/migration-check/inputs/examples.json","/samples/migration-check/inputs/prompt.json","/samples/migration-check/inputs/settings.json","/samples/migration-check/inputs/task.json"],"bytes":3510,"checks":["all 10 examples were scored","at least 9 of 10 outputs parse against the schema","no model call failed","at least 7 of 10 outputs agree with the reference on every field","the verdict is one of go, no-go or inconclusive","the sealed record verifies","every stated number recomputes from the per-example entries","the record was issued by this server","a record with one example's pass mark changed no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: 20 fictional support tickets for a made-up shop (the first 10 here). The reference outputs were written for this demo, not produced by a closed API. Written for the Decosa demo, 25 Sep 2026. Part of decosa-api, AGPL-3.0-or-later.","about":"Ten fictional support tickets for a made-up homeware shop, each with the JSON a hand-written reference gave (category, priority, refund_requested, order_id). The open model answers the same prompt; its outputs must parse against the schema and mostly agree field by field, and the sealed record must verify, recompute every stated number and fail once one example's pass mark is changed.","run":{"containers":"docker compose exec api python scripts/rehearse.py migration-check","checkout":"python scripts/rehearse.py migration-check --bundle migration-check.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py migration-check"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=migration-check","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":38,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":192,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1377.6,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/migration-check","json":"/use-cases/migration-check.json","metrics":"/metrics/migration-check","console":"/tools/developer/migration-check","console_sample":"/tools/developer/migration-check?sample=1&autorun=0","stack":"/tools/developer/migration-check#stack","try_live":"/tools/developer/migration-check","watch":"/tools/developer/migration-check","build":"/tools/developer/migration-check#build","self_host":"/tools/developer/migration-check#self-host","prompts":{"hosted":"/prompts/migration-check-hosted.md","selfhost":"/prompts/migration-check-selfhost.md","assemble":"/prompts/migration-check-assemble.md","mac":"/prompts/migration-check-mac.md"}}}