Skip to content
decosa

09 · Software and AI ops · Compliance and trust · live

Endpoint auditor

Open the toolJSON

Eval results

Not held outRun 24 Sep 2026

  • Swap caught: Qwen3.5-4B-Base served as qwen3.8-27bfail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergencessyntheticreport aud_7be2d97ae80e20abdf34
  • Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963syntheticreport aud_8a7c979db05198f582de
  • Reference noise, Qwen3.8-27B stack (6 runs)worst greedy repeat 5/10; mean |Δ logprob| ≤ 0.040; top-5 overlap ≥ 0.872syntheticn = 6fixture qwen3.8-27b.json
  • Hosted gateway route vs referencepass; greedy 7/10 (band ≥ 4/10), canaries 8/10 = referencesyntheticreport aud_687900df0710578a42a8
  • Raw engine route vs referencepass; greedy 10/10, logprobs identical over 244 tokenssyntheticreport aud_b5a3508ddb34527f490e
  • Claim of BF16 weights ("Qwen/Qwen3.8-27B") when only the NVFP4 reference existsinconclusive, 4 of 4 claim runs (before 28 Sep it could be signed pass)syntheticn = 4docs/evals/auditor-claims.md; our own direct engine
  • Genuine endpoint, correct claim, 28 Sep re-runhosted pass 3/3; direct engine pass 4/6 (2 drift in one run)syntheticn = 9right after the production model restart; the noise band is too tight for a busy card

Dataset

Audit reports and reference fixtures run on our server: two deliberate swaps (a smaller model and a re-quantised model under a false name), six reference runs to measure noise, and the hosted and raw routes against the reference.

Caveats

  • Only two planted swaps, both set up by the builder; subtler substitutions were not tested.
  • A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.
  • On a busy card, a genuine self-hosted endpoint was flagged drift in about one run in three (1 of 3 on 24 Sep, 2 of 6 on 28 Sep), even though the re-check agreed on 28 Sep. Treat a single drift as a prompt to re-run, not a finding.
  • DeepSeek-V4-Flash reference fixture (best tier) not measured yet.
  • The claimed precision decides the reference: an official name such as Qwen/Qwen3.8-27B means BF16, and with no BF16 reference the verdict is inconclusive, not pass (28 Sep 2026).

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
40 s
Receipts
22
Model calls
n/a
Tokens
n/a
Cost per run
$0.002

Self-host verification

Verified on 25 Sep 2026: Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, audit of a local Qwen3.8-27B vLLM

Verified on 2026-09-25: signing key created, both reference fixtures listed, a full audit of a local Qwen3.8-27B vLLM (equivalent to the documented target) returned pass in 29 s with 23 probes, the report verified with the documented Python snippet and failed after an edit. The prompt's ./keys bind mount is not writable by the image user; the key was kept in the data volume instead (fix in progress).

Rehearsal bundle: auditor.zip (16 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted figures are for the short audit (context probe off, 22 probes). The console's default run adds a ~12k-token context probe.
  • The hosted auditor only reaches public https:// endpoints; audit private or internal endpoints with the self-hosted auditor.
  • On a busy card, one self-hosted run in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it.
  • When the gateway is slow, the console's availability check marks the hosted target as not running and plays its recorded audit instead (seen 2026-09-25); POST /audit/runs still ran live.
  • A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.
  • Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Probe runner, scorer and signer (no model; runs on CPU)decosa-api auditor (decosa_api/verticals/auditor)AGPL-3.0-or-later
  • Reference model (golden outputs, logprobs, noise band)Qwen3.8-27B (NVFP4)Apache-2.0
  • Small reference model (swap and quantisation demos)Qwen3.5-4B-Base (BF16)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · audits only, no GPU (2)
  • Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b: fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergencesreport aud_7be2d97ae80e20abdf34, our server 2026-09-24
  • Quantisation drift caught: FP8 re-quant claimed as BF16 (4B): drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963report aud_8a7c979db05198f582de, our server 2026-09-24
Standard · one 96 GB card (hosted demo) (4)
  • Reference noise, Qwen3.8-27B stack (16 runs, 12 of them under load): worst greedy repeat 4/10; mean |Δ logprob| ≤ 0.055; top-5 overlap ≥ 0.826fixture qwen3.8-27b.json, our server 2026-09-30 (re-recorded after the server was restarted with image and video input on 28 Sep)
  • Hosted gateway route vs reference: pass; greedy at or above the band (≥ 3/10), canaries equal to the reference. The nightly check runs this routethe nightly check (its latest result is under 'How we tested it')
  • Raw engine route vs reference: pass in 10 of 10 audits on 30 Sep 2026; greedy 4 to 7 of 10 (band ≥ 3/10), top-5 overlap 0.835 to 0.893 (band ≥ 0.806)docs/evals/auditor-claims.md, 30 Sep 2026
  • Engine drift on the same weights (coding benchmark, /500): 478 NVFP4 + MTP; 387 FP8 eager; 193 FP8 + MTP; 62-63 llama.cpp CUDAcoding-agent-bench README (not re-run by the auditor)
Best · two 96 GB cards (1)
  • DeepSeek-V4-Flash reference fixture: not measured yetnot measured yet
Wanted · references for the most-used open models (1)
  • substitution detection on these models: not measured yet

How we measure · All tools