{"schema_version":"1","site":"https://decosa.ai","id":"auditor","num":"09","name":"Endpoint auditor","tool_name":"Check your provider serves the model you pay for","short":"Auditor","blurb":"Checks whether an OpenAI-compatible endpoint really serves the model it claims, at the quality it claims, and signs the result.","status":"live","labels":{"industry":["software","compliance-trust"],"job":["attest","review"],"input":["endpoint"],"deploy":["hosted","selfhost"],"status":"live","output":["record"],"data":["none"],"hardware":"cpu","licence":"permissive"},"industries":["software","compliance-trust"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["endpoint-audit","signed-record"],"models":"Reference runs of Qwen3.8-27B · Qwen3.5-4B","where":"Hosted or self-host","hardware":"No GPU needed to audit. References are recorded on 1× RTX PRO 6000 (96 GB).","final_artifact":"A signed audit report.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Signed report, receipted probes","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":39822,"p95_ms":null,"runs":null,"receipts_per_run":22,"cost_per_run_usd":0.0018},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, audit of a local Qwen3.8-27B vLLM","notes":"Verified on 2026-09-25: signing key created, both reference fixtures listed, a full audit of a local Qwen3.8-27B vLLM (equivalent to the documented target) returned pass in 29 s with 23 probes, the report verified with the documented Python snippet and failed after an edit. The prompt's ./keys bind mount is not writable by the image user; the key was kept in the data volume instead (fix in progress)."},"known_limits":["Hosted figures are for the short audit (context probe off, 22 probes). The console's default run adds a ~12k-token context probe.","The hosted auditor only reaches public https:// endpoints; audit private or internal endpoints with the self-hosted auditor.","On a busy card, one self-hosted run in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it.","When the gateway is slow, the console's availability check marks the hosted target as not running and plays its recorded audit instead (seen 2026-09-25); POST /audit/runs still ran live.","A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.","Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b","value":"fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences","unit":null,"n":null,"split":"synthetic","note":"report aud_7be2d97ae80e20abdf34"},{"name":"Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)","value":"drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963","unit":null,"n":null,"split":"synthetic","note":"report aud_8a7c979db05198f582de"},{"name":"Reference noise, Qwen3.8-27B stack (6 runs)","value":"worst greedy repeat 5/10; mean |Δ logprob| ≤ 0.040; top-5 overlap ≥ 0.872","unit":null,"n":6,"split":"synthetic","note":"fixture qwen3.8-27b.json"},{"name":"Hosted gateway route vs reference","value":"pass; greedy 7/10 (band ≥ 4/10), canaries 8/10 = reference","unit":null,"n":null,"split":"synthetic","note":"report aud_687900df0710578a42a8"},{"name":"Raw engine route vs reference","value":"pass; greedy 10/10, logprobs identical over 244 tokens","unit":null,"n":null,"split":"synthetic","note":"report aud_b5a3508ddb34527f490e"},{"name":"Claim of BF16 weights (\"Qwen/Qwen3.8-27B\") when only the NVFP4 reference exists","value":"inconclusive, 4 of 4 claim runs (before 28 Sep it could be signed pass)","unit":null,"n":4,"split":"synthetic","note":"docs/evals/auditor-claims.md; our own direct engine"},{"name":"Genuine endpoint, correct claim, 28 Sep re-run","value":"hosted pass 3/3; direct engine pass 4/6 (2 drift in one run)","unit":null,"n":9,"split":"synthetic","note":"right after the production model restart; the noise band is too tight for a busy card"}],"dataset":"Audit reports and reference fixtures run on our server: two deliberate swaps (a smaller model and a re-quantised model under a false name), six reference runs to measure noise, and the hosted and raw routes against the reference.","held_out":false,"caveats":["Only two planted swaps, both set up by the builder; subtler substitutions were not tested.","A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.","On a busy card, a genuine self-hosted endpoint was flagged drift in about one run in three (1 of 3 on 24 Sep, 2 of 6 on 28 Sep), even though the re-check agreed on 28 Sep. Treat a single drift as a prompt to re-run, not a finding.","DeepSeek-V4-Flash reference fixture (best tier) not measured yet.","The claimed precision decides the reference: an official name such as Qwen/Qwen3.8-27B means BF16, and with no BF16 reference the verdict is inconclusive, not pass (28 Sep 2026)."],"date":"2026-09-24","doc_url":null},"stack":{"summary":"The auditor sends a fixed suite of probes at temperature 0 to any OpenAI-compatible endpoint: ten golden prompts compared with a reference run of the claimed weights (text, engine token counts and, where the endpoint exposes them, per-token logprobs), ten graded canaries, a ~12k-token context probe and a streamed throughput probe. It classifies the endpoint as pass, drift or fail, re-runs the probes before any drift or fail is signed, and signs the report with a dedicated Ed25519 key that anyone can verify. It is for teams buying open-model inference, routers and model labs who want evidence, not a provider's word, that the weights and engine behind an endpoint are the ones they pay for.","tagline":"Continuous, signed checks that an OpenAI-compatible endpoint serves the model and quality it claims.","deployment":"hosted-or-self-host","regulatory_note":"Auditing a third-party endpoint uses your own key and credits with that provider: the hosted auditor holds the key in memory for one run and never stores or logs it. Check your provider's terms before publishing results; reports are private by default. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3.5-4B-Base).","components":[{"id":"auditor","role":"Probe runner, scorer and signer (no model; runs on CPU)","name":"decosa-api auditor (decosa_api/verticals/auditor)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12, httpx; suite auditor-suite-v1","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard","best"],"alternative_to":null},{"id":"ref-27b","role":"Reference model (golden outputs, logprobs, noise band)","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":null,"tiers":["standard","best"],"alternative_to":null},{"id":"ref-4b","role":"Small reference model (swap and quantisation demos)","name":"Qwen3.5-4B-Base (BF16)","hf_repo":"Qwen/Qwen3.5-4B-Base","license":"Apache-2.0","params":"4B","quant":"BF16 (reference); FP8 on the fly for the drift demo","vram_gb":13,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, --enforce-eager, max-model-len 16384","receipt_coverage":"none","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"ref-dsv4","role":"Larger reference (planned)","name":"DeepSeek-V4-Flash (NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts (MTP draft experts MXFP4)","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM, B12X native sm_120 build, TP2 (see the code use case)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["best"],"alternative_to":null},{"id":"glm-wanted","role":"Reference for GLM-5.3-Flash endpoints","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null},{"id":"dsv41-wanted","role":"Reference for DeepSeek-V4.1-Flash endpoints","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · audits only, no GPU","summary":"Run the auditor against any endpoint with the signed reference fixtures we ship. You cannot record new references.","components":["auditor","ref-4b"],"hardware":"Any Linux or macOS machine with Python 3.11+","quality_evidence":[{"metric":"Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b","value":"fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences","source":"report aud_7be2d97ae80e20abdf34, our server 2026-09-24"},{"metric":"Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)","value":"drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963","source":"report aud_8a7c979db05198f582de, our server 2026-09-24"}],"latency_note":"measured: seconds per audit, depending on the endpoint and whether a re-check runs.","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":"Signed report; probes to non-the network endpoints carry no receipts.","hosting":null},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","summary":"Adds recording your own references on the pinned Qwen3.8-27B stack, with its noise band measured idle and under load.","components":["auditor","ref-27b","ref-4b"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Reference noise, Qwen3.8-27B stack (16 runs, 12 of them under load)","value":"worst greedy repeat 4/10; mean |Δ logprob| ≤ 0.055; top-5 overlap ≥ 0.826","source":"fixture qwen3.8-27b.json, our server 2026-09-30 (re-recorded after the server was restarted with image and video input on 28 Sep)"},{"metric":"Hosted gateway route vs reference","value":"pass; greedy at or above the band (≥ 3/10), canaries equal to the reference. The nightly check runs this route","source":"the nightly check (its latest result is under 'How we tested it')"},{"metric":"Raw engine route vs reference","value":"pass in 10 of 10 audits on 30 Sep 2026; greedy 4 to 7 of 10 (band ≥ 3/10), top-5 overlap 0.835 to 0.893 (band ≥ 0.806)","source":"docs/evals/auditor-claims.md, 30 Sep 2026"},{"metric":"Engine drift on the same weights (coding benchmark, /500)","value":"478 NVFP4 + MTP; 387 FP8 eager; 193 FP8 + MTP; 62-63 llama.cpp CUDA","source":"coding-agent-bench README (not re-run by the auditor)"}],"latency_note":"measured: seconds per audit on the raw route, about twice that on the receipted gateway route.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every probe to the hosted route has a gateway-signed receipt; the report lists their ids.","hosting":null},{"id":"best","label":"Best · two 96 GB cards","summary":"Adds a DeepSeek-V4-Flash reference so endpoints selling that model can be audited.","components":["auditor","ref-27b","ref-dsv4"],"hardware":"2x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"DeepSeek-V4-Flash reference fixture","value":"not measured yet","source":"not measured yet"}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"DeepSeek-V4-Flash is not a hosted model, so its reference calls would carry no receipts.","hosting":null},{"id":"wanted","label":"Wanted · references for the most-used open models","summary":"Reference runs of GLM-5.3-Flash and DeepSeek-V4.1-Flash, the two most-used open models, so endpoints that sell them can be audited against the real thing. Neither fits the reference box today. Not served yet.","components":["auditor","glm-wanted","dsv41-wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). Estimate.","quality_evidence":[{"metric":"substitution detection on these models","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /audit/targets, POST /audit/runs (SSE), GET /audit/reports/{id}, GET /audit/signing-key, GET /audit/references, POST /audit/verify."},{"name":"auditor-swap-4b (demo only)","port":8201,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Serves Qwen3.5-4B-Base with FP8 on-the-fly quantisation under two names: qwen3.8-27b (the swap) and qwen3.5-4b-base (the drift). Loopback only, about 13 GB on GPU0."}],"tools":[{"name":"verify_audit_report.py","url":null,"license":"Apache-2.0","purpose":"Offline check of a signed report with only the cryptography package (decosa-api scripts/)."},{"name":"coding-agent-bench","url":null,"license":null,"purpose":"The owner's coding benchmark and source of the engine-drift evidence: the same Qwen3.8-27B weights scored 478, 387, 193 and 62-63 out of 500 depending on engine."}],"hardware":[{"tier":"Any CPU, no GPU","fits":true,"notes":"Running audits. Reference fixtures ship with the auditor, signed."},{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Recording the Qwen3.8-27B reference on the pinned stack; measured on our server."},{"tier":"2x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Needed to record a DeepSeek-V4-Flash reference. Not done yet."}],"latency":[{"lane":"full audit, raw engine route with logprobs (23 requests, no re-check)","typical_ms":7686,"source":"measured on our server 2026-09-24 (report aud_b5a3508ddb34527f490e)"},{"lane":"full audit, hosted gateway route with receipts (23 requests)","typical_ms":15335,"source":"measured on our server 2026-09-24 (report aud_687900df0710578a42a8)"},{"lane":"full audit with re-check, 4B endpoint (43 requests)","typical_ms":21745,"source":"measured on our server 2026-09-24 (report aud_7be2d97ae80e20abdf34)"},{"lane":"reference decode speed, Qwen3.8-27B NVFP4 + MTP3, single stream","typical_ms":8,"source":"measured on our server 2026-09-24: 129.5 tok/s median of 3 (fixture qwen3.8-27b)"}],"benchmark":null,"notes":["The gateway route reports its own metered token estimates (for example 34 prompt tokens where the engine counts 26 for the same request), so the token-count fingerprint is skipped on that route; greedy match carries identity there.","Greedy decoding on the pinned 27B stack is not batch-invariant: under concurrent load only 5 of 10 golden texts reproduced and the mean |Δ logprob| reached 0.040. The pass band is set from that measured spread (6 runs: 3 idle, 3 with 4 concurrent streams), not from an idle run.","An FP8 re-quantisation of a 4B model moves logprobs by about as much as batch noise does on the 27B. It is caught through the top-5 overlap and greedy match against the 4B's own tighter band, not by the mean shift alone. Small quantisation changes on a busy 27B endpoint may stay inside its band: the auditor would say pass, and that limit is real.","Engine-drift evidence we did not re-run here: the same Qwen3.8-27B weights scored 478/500 on vLLM NVFP4 + MTP, 387 on vLLM FP8 eager, 193 on vLLM FP8 + MTP and 62-63 on llama.cpp CUDA (coding-agent-bench README). A second 27B engine config could not be stood up without stopping a live service."]},"buyer_facts":[{"label":"Your provider's API key","value":"Held in memory for one run only: never logged, stored or written into the report. The probes bill your provider (about 23-43 short requests plus the context probe)."},{"label":"Data retention","value":"Signed reports are stored on the server and readable by anyone with the report id; a report of your endpoint records its URL and model name, never the key."},{"label":"Hardware","value":"CPU only: comparison data ships as signed reference fixtures. Recording new fixtures needs the model's pinned GPU stack."},{"label":"Typical hosted cost","value":"A fraction of a cent of gateway time for a short audit of our own hosted route (measured). Each run shows its own measured cost."}],"data_handling":{"page":"/data#auditor","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Signed reports are stored on the server and readable by anyone with the report id; a report of your endpoint records its URL and model name, never the key.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"The endpoint you audit","route":"both","sends":"your-system","what":"About 23-43 fixed probe prompts (not your data) and the provider API key you give, held in memory for one run. The probes bill your provider.","default":"always","off":null}]},"console":{"href":"/tools/developer/auditor","input":"endpoint","lanes":[{"id":"identity","title":"Identity checks","kind":"list"},{"id":"quality","title":"Quality canaries","kind":"list"},{"id":"performance","title":"Latency and throughput","kind":"list"},{"id":"verdict","title":"Signed verdict","kind":"json"}],"samples":[{"n":1,"id":"auditor-hosted","title":"Hosted","deep_link":"/tools/developer/auditor?sample=1&autorun=0"},{"n":2,"id":"auditor-direct","title":"Direct","deep_link":"/tools/developer/auditor?sample=2&autorun=0"},{"n":3,"id":"auditor-swap","title":"Swap","deep_link":"/tools/developer/auditor?sample=3&autorun=0"},{"n":4,"id":"auditor-quant","title":"Quant","deep_link":"/tools/developer/auditor?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/auditor-hosted.md","selfhost":"/prompts/auditor-selfhost.md","assemble":"/prompts/auditor-assemble.md","mac":"/prompts/auditor-mac.md"},"rehearsal":{"bundle":"/samples/auditor.zip","bundle_url":"https://decosa.ai/samples/auditor.zip","folder":"/samples/auditor/","expected":"/samples/auditor/expected.json","files":["/samples/auditor/expected.json","/samples/auditor/inputs/own-endpoint.example.json","/samples/auditor/inputs/qwen3.8-27b.json","/samples/auditor/inputs/qwen3.8-27b.sig"],"bytes":16155,"checks":["at least one audit target is up on this server","the audit finishes without errors","the audit compares against the shipped reference fixture","no probe call failed","greedy outputs match the reference within its band","the canary score is within the reference band","the verdict is pass (or inconclusive after a re-check near the band edge)","the signed report verifies against this auditor's key","a report with one check value changed no longer verifies","every model call has a signed receipt"],"licence":"The probe suite and reference fixture are part of decosa-api, AGPL-3.0-or-later (the fixture was recorded from Qwen3.8-27B NVFP4, Apache-2.0 weights). No user data: every probe is a synthetic prompt.","about":"A short audit (22 probes, long-context check off) of the first available target on this server (GET /audit/targets: the hosted gateway route, or on a self-host box its local model), compared with the signed reference fixture for Qwen3.8-27B that ships here. The endpoint must match the reference on identity and quality, and the signed report must verify against this auditor's key and fail once a check value is changed. The route takes a target id or your own endpoint (inputs/own-endpoint.example.json), not a raw fixture: the fixture is shipped for reading.","run":{"containers":"docker compose exec api python scripts/rehearse.py auditor","checkout":"python scripts/rehearse.py auditor --bundle auditor.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py auditor"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=auditor","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":13,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":70.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":249.6,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1320,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":16}},"links":{"page":"/tools/developer/auditor","json":"/use-cases/auditor.json","metrics":"/metrics/auditor","console":"/tools/developer/auditor","console_sample":"/tools/developer/auditor?sample=1&autorun=0","stack":"/tools/developer/auditor#stack","try_live":"/tools/developer/auditor","watch":"/tools/developer/auditor","build":"/tools/developer/auditor#build","self_host":"/tools/developer/auditor#self-host","prompts":{"hosted":"/prompts/auditor-hosted.md","selfhost":"/prompts/auditor-selfhost.md","assemble":"/prompts/auditor-assemble.md","mac":"/prompts/auditor-mac.md"}}}