09 · Software and AI ops · Compliance and trust · live
Endpoint auditor
Eval results
Not held outRun 24 Sep 2026
- Swap caught: Qwen3.5-4B-Base served as qwen3.8-27bfail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergencessyntheticreport aud_7be2d97ae80e20abdf34
- Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963syntheticreport aud_8a7c979db05198f582de
- Reference noise, Qwen3.8-27B stack (6 runs)worst greedy repeat 5/10; mean |Δ logprob| ≤ 0.040; top-5 overlap ≥ 0.872syntheticn = 6fixture qwen3.8-27b.json
- Hosted gateway route vs referencepass; greedy 7/10 (band ≥ 4/10), canaries 8/10 = referencesyntheticreport aud_687900df0710578a42a8
- Raw engine route vs referencepass; greedy 10/10, logprobs identical over 244 tokenssyntheticreport aud_b5a3508ddb34527f490e
- Claim of BF16 weights ("Qwen/Qwen3.8-27B") when only the NVFP4 reference existsinconclusive, 4 of 4 claim runs (before 28 Sep it could be signed pass)syntheticn = 4docs/evals/auditor-claims.md; our own direct engine
- Genuine endpoint, correct claim, 28 Sep re-runhosted pass 3/3; direct engine pass 4/6 (2 drift in one run)syntheticn = 9right after the production model restart; the noise band is too tight for a busy card
Dataset
Audit reports and reference fixtures run on our server: two deliberate swaps (a smaller model and a re-quantised model under a false name), six reference runs to measure noise, and the hosted and raw routes against the reference.
Caveats
- Only two planted swaps, both set up by the builder; subtler substitutions were not tested.
- A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.
- On a busy card, a genuine self-hosted endpoint was flagged drift in about one run in three (1 of 3 on 24 Sep, 2 of 6 on 28 Sep), even though the re-check agreed on 28 Sep. Treat a single drift as a prompt to re-run, not a finding.
- DeepSeek-V4-Flash reference fixture (best tier) not measured yet.
- The claimed precision decides the reference: an official name such as Qwen/Qwen3.8-27B means BF16, and with no BF16 reference the verdict is inconclusive, not pass (28 Sep 2026).
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 40 s
- Receipts
- 22
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 25 Sep 2026: Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, audit of a local Qwen3.8-27B vLLM
Verified on 2026-09-25: signing key created, both reference fixtures listed, a full audit of a local Qwen3.8-27B vLLM (equivalent to the documented target) returned pass in 29 s with 23 probes, the report verified with the documented Python snippet and failed after an edit. The prompt's ./keys bind mount is not writable by the image user; the key was kept in the data volume instead (fix in progress).
Rehearsal bundle: auditor.zip (16 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted figures are for the short audit (context probe off, 22 probes). The console's default run adds a ~12k-token context probe.
- The hosted auditor only reaches public https:// endpoints; audit private or internal endpoints with the self-hosted auditor.
- On a busy card, one self-hosted run in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it.
- When the gateway is slow, the console's availability check marks the hosted target as not running and plays its recorded audit instead (seen 2026-09-25); POST /audit/runs still ran live.
- A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.
- Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Probe runner, scorer and signer (no model; runs on CPU)decosa-api auditor (decosa_api/verticals/auditor)AGPL-3.0-or-later
- Reference model (golden outputs, logprobs, noise band)Qwen3.8-27B (NVFP4)Apache-2.0
- Small reference model (swap and quantisation demos)Qwen3.5-4B-Base (BF16)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · audits only, no GPU (2)
- Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b: fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergencesreport aud_7be2d97ae80e20abdf34, our server 2026-09-24
- Quantisation drift caught: FP8 re-quant claimed as BF16 (4B): drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963report aud_8a7c979db05198f582de, our server 2026-09-24
Standard · one 96 GB card (hosted demo) (4)
- Reference noise, Qwen3.8-27B stack (16 runs, 12 of them under load): worst greedy repeat 4/10; mean |Δ logprob| ≤ 0.055; top-5 overlap ≥ 0.826fixture qwen3.8-27b.json, our server 2026-09-30 (re-recorded after the server was restarted with image and video input on 28 Sep)
- Hosted gateway route vs reference: pass; greedy at or above the band (≥ 3/10), canaries equal to the reference. The nightly check runs this routethe nightly check (its latest result is under 'How we tested it')
- Raw engine route vs reference: pass in 10 of 10 audits on 30 Sep 2026; greedy 4 to 7 of 10 (band ≥ 3/10), top-5 overlap 0.835 to 0.893 (band ≥ 0.806)docs/evals/auditor-claims.md, 30 Sep 2026
- Engine drift on the same weights (coding benchmark, /500): 478 NVFP4 + MTP; 387 FP8 eager; 193 FP8 + MTP; 62-63 llama.cpp CUDAcoding-agent-bench README (not re-run by the auditor)
Best · two 96 GB cards (1)
- DeepSeek-V4-Flash reference fixture: not measured yetnot measured yet
Wanted · references for the most-used open models (1)
- substitution detection on these models: not measured yet