Skip to content
decosa
LiveHostedSelf-hostMac

Check your provider serves the model you pay for

Find out whether an OpenAI-style endpoint really serves the model it claims, or a smaller or more compressed one. You get a signed pass, drift or fail report you can put in the vendor file, and anyone can re-check it.

On production40 smedian on production (2026-09-25); slower when the service is busy
List price~$0.18 per 100 auditsmeasured, at list price

Built on: Endpoint audit, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the endpoint auditor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on No GPU needed to audit. References are recorded on 1× RTX PRO 6000 (96 GB)..
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
auditor

Use the hosted API

# Decosa Endpoint auditor: use the hosted API

You are wiring Decosa's endpoint auditor into this project. It checks whether an OpenAI-compatible endpoint really
serves the model it claims, at the quality it claims, and returns a report signed with the auditor's Ed25519 key.
Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "auditor"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz). Over a limit: HTTP 429 with `Retry-After`.
3. One audit at a time per demo token. An audit needs at least 3,000 generated tokens of budget left (402 otherwise).

## Endpoints (no token needed unless marked)
- `GET /audit/signing-key` → `{"alg": "ed25519", "pubkey", "key_id", "domain", "canonical"}`. Pin this key.
- `GET /audit/targets` → our own demo endpoints `[{id, label, claimed_model, kind, note, expect, available}]`.
- `GET /audit/references` → reference fixtures `[{id, aliases, model, engine, recorded_at, fixture_sha256, fixture_sig, noise}]`.
- `POST /audit/runs` (token) → SSE. Body is either `{"target": "<demo id>"}` or
  `{"base_url": "https://…/v1", "model": "<name on that endpoint>", "claimed_model"?: "<reference alias>", "api_key"?: "<their key>", "context_tokens"?: 0-32000}`.
  - Hosted audits only reach public `https://` endpoints (400 for private, loopback or plain-http URLs).
  - `api_key` is the key for the audited provider. It is held in memory for this run only: never logged, stored or put in the report. The probes bill that provider (about 23-43 short requests plus the context probe).
- `GET /audit/reports/{id}` → a signed report. `GET /audit/reports` lists recent reports for our demo targets only.
- `POST /audit/verify` with `{"report": {...}}` → `{"valid_signature", "signed_by_this_auditor", "auditor_pubkey"}` (a convenience; verify locally too).

## SSE events from `/audit/runs`
- `{"type": "ready", "target", "claimed_model", "reference", "suite": {"id", "sha256", "goldens", "canaries"}}`
- `{"type": "lane", "lane": "identity" | "quality" | "performance", "data": {"checks": [{"id", "title", "status": "pass"|"warn"|"fail"|"skip", "value", "band", "detail"}], "progress"}}`. The latest event per lane is its full state: replace, don't append.
- `{"type": "receipt", ...}` for each probe that went through the Decosa API (hosted target only).
- `{"type": "lane", "lane": "verdict", "data": {"status", "reasons", "report", "report_url"}}`, then `{"type": "done", "summary": {"verdict", "report", "report_url"}}`, then a `budget` event.
- `{"type": "error", "message"}` if the run stops.

## Verdicts
`pass` (inside the reference stack's measured spread), `drift` (same family, different numerics or quality),
`fail` (not the claimed weights), `inconclusive` (too many probe errors, no reference, or a re-check that disagreed).
Drift and fail are only signed after a second pass agrees.

## Verify a report (Python, `pip install cryptography`)
```python
import json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
API = "https://api.decosa.ai"
report = json.load(urllib.request.urlopen(f"{API}/audit/reports/<id>"))
pub = json.load(urllib.request.urlopen(f"{API}/audit/signing-key"))["pubkey"]
assert report["signer"]["pubkey"] == pub
body = {k: v for k, v in report.items() if k != "signature"}
msg = b"decosa.audit.report.v1\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(report["signature"]), msg)
```
In a browser: rebuild the same bytes with a sorted-keys `JSON.stringify` and check them with WebCrypto `Ed25519`.

## Example: run an audit and print the verdict (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"base_url": "https://api.example.com/v1", "model": "qwen3.8-27b", "api_key": os.environ["PROVIDER_KEY"], "context_tokens": 4000}
with httpx.stream("POST", f"{API}/audit/runs", json=body, headers=H, timeout=600) as r:
    if r.status_code != 200:
        r.read()
        raise SystemExit(f"{r.status_code}: {r.json().get('error') or r.json()}")   # e.g. a private URL or a host that does not resolve
    for line in r.iter_lines():
        if line.startswith("data: "):
            ev = json.loads(line[6:])
            if ev["type"] == "done":
                print(ev["summary"]["verdict"], ev["summary"]["report_url"])
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Endpoint auditor: run it yourself (containers)

You are setting up the Decosa endpoint auditor to run on this machine, so the API keys of the endpoints I audit, and
the endpoints themselves, never leave my network. It needs no GPU. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/auditor.zip (16 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py auditor` (the api image carries the same bundle under /app/rehearsal/auditor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py auditor --bundle auditor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least one audit target is up on this server", "the audit finishes without errors", "the audit compares against the shipped reference fixture"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install).
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Set these for the `api` service: `DECOSA_AUDITOR_KEY=/data/keys/auditor_ed25519.key`,
   `DECOSA_AUDITOR_CREATE_KEY=1` (the first start creates the key, mode 0600) and `DECOSA_AUDIT_ALLOW_PRIVATE=1`
   (so it may audit endpoints on this network). Keep `/data` on the named volume the compose file declares: the key
   is created inside it (a host `./keys` folder is not writable by the api's user, uid 10001). Back it up with
   `docker compose cp api:/data/keys ./auditor-keys-backup` and never commit it.
3. Pull and start: `docker compose pull && docker compose up -d`.
4. Check: `curl -fsS http://localhost:<PORT>/audit/signing-key` returns the auditor's public key. Show it to me: it is what
   others pin to verify my reports. `GET /audit/references` lists the signed reference fixtures.
5. Smoke test: get a token with `POST /demo/session {"vertical":"auditor"}`, then
   `POST /audit/runs` with `{"base_url": "<endpoint>/v1", "model": "<name>", "claimed_model": "qwen3.8-27b"}` and read the
   SSE stream until `done`. Pass a provider key as `api_key` from an environment variable; never print it.
6. Report back: the public key and key id, the references listed, and the smoke-test verdict and report id.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 16 GB of unified memory or more): use https://decosa.ai/prompts/auditor-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 33 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 41 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (70.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (70.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (16 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (16 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"auditor"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py auditor

Download the mock-data bundle (16 KB, 10 checks)expected.json

A short audit (22 probes, long-context check off) of the first available target on this server (GET /audit/targets: the hosted gateway route, or on a self-host box its local model), compared with the signed reference fixture for Qwen3.8-27B that ships here. The endpoint must match the reference on identity and quality, and the signed report must verify against this auditor's key and fail once a check value is changed. The route takes a target id or your own endpoint (inputs/own-endpoint.example.json), not a raw fixture: the fixture is shipped for reading.

What the rehearsal checks
  • at least one audit target is up on this server
  • the audit finishes without errors
  • the audit compares against the shipped reference fixture
  • no probe call failed
  • greedy outputs match the reference within its band
  • the canary score is within the reference band
  • the verdict is pass (or inconclusive after a re-check near the band edge)
  • the signed report verifies against this auditor's key
  • a report with one check value changed no longer verifies
  • every model call has a signed receipt

Licence: The probe suite and reference fixture are part of decosa-api, AGPL-3.0-or-later (the fixture was recorded from Qwen3.8-27B NVFP4, Apache-2.0 weights). No user data: every probe is a synthetic prompt.

Prompt for your coding agent

# Decosa Endpoint auditor: run it yourself (containers)

You are setting up the Decosa endpoint auditor to run on this machine, so the API keys of the endpoints I audit, and
the endpoints themselves, never leave my network. It needs no GPU. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/auditor.zip (16 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py auditor` (the api image carries the same bundle under /app/rehearsal/auditor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py auditor --bundle auditor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least one audit target is up on this server", "the audit finishes without errors", "the audit compares against the shipped reference fixture"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install).
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Set these for the `api` service: `DECOSA_AUDITOR_KEY=/data/keys/auditor_ed25519.key`,
   `DECOSA_AUDITOR_CREATE_KEY=1` (the first start creates the key, mode 0600) and `DECOSA_AUDIT_ALLOW_PRIVATE=1`
   (so it may audit endpoints on this network). Keep `/data` on the named volume the compose file declares: the key
   is created inside it (a host `./keys` folder is not writable by the api's user, uid 10001). Back it up with
   `docker compose cp api:/data/keys ./auditor-keys-backup` and never commit it.
3. Pull and start: `docker compose pull && docker compose up -d`.
4. Check: `curl -fsS http://localhost:<PORT>/audit/signing-key` returns the auditor's public key. Show it to me: it is what
   others pin to verify my reports. `GET /audit/references` lists the signed reference fixtures.
5. Smoke test: get a token with `POST /demo/session {"vertical":"auditor"}`, then
   `POST /audit/runs` with `{"base_url": "<endpoint>/v1", "model": "<name>", "claimed_model": "qwen3.8-27b"}` and read the
   SSE stream until `done`. Pass a provider key as `api_key` from an environment variable; never print it.
6. Report back: the public key and key id, the references listed, and the smoke-test verdict and report id.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 16 GB of unified memory or more): use https://decosa.ai/prompts/auditor-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierEndpoint auditor on GeForce RTX 5090: use the Lite · audits only, no GPU tier

The standard tier does not fit: Needs about 41 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.

Lite · audits only, no GPU: what changes

Nothing: it runs as listed in the stack.

Memory per component
  • Probe runner, scorer and signer: decosa-api auditor (decosa_api/verticals/auditor). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Small reference model: Qwen3.5-4B-Base (BF16). ~13 GB (from stack.json). vram_gb 13 in stack.json.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Endpoint auditor, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Endpoint auditor on my hardware

Fetch https://decosa.ai/prompts/auditor-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=auditor)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · audits only, no GPU (lite). Fit check: runs, about 13 GB of 32 GB used.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Probe runner, scorer and signer: decosa-api auditor (decosa_api/verticals/auditor), CPU
- Small reference model: Qwen3.5-4B-Base (BF16) (Qwen/Qwen3.5-4B-Base), 13 GB

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.5-4B-Base (BF16) ~13 GB (41%); about 19 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/auditor-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 16 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Endpoint auditor: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Endpoint auditor on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 16 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/auditor.zip (16 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py auditor` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least one audit target is up on this server", "the audit finishes without errors", "the audit compares against the shipped reference fixture"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Probe runner, scorer and signer (no model; runs on CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Reference model (golden outputs, logprobs, noise band) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Small reference model (swap and quantisation demos) | BF16 on vLLM | MLX (mlx_lm.server) | Runs, not measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 16 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py auditor`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Auditing an endpoint needs no model: the suite and the signed reference fixtures run on CPU. Auditing this Mac's own MLX model against the NVFP4 reference fails, as it should: they are different weights.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 40 s · ~$0.002 per run · 22 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, audit of a local Qwen3.8-27B vLLM

Measured cost to run: about $0.18 per 100 audits (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: signing key created, both reference fixtures listed, a full audit of a local Qwen3.8-27B vLLM (equivalent to the documented target) returned pass in 29 s with 23 probes, the report verified with the documented Python snippet and failed after an edit. The prompt's ./keys bind mount is not writable by the image user; the key was kept in the data volume instead (fix in progress).

Known limits (6)
  • Hosted figures are for the short audit (context probe off, 22 probes). The console's default run adds a ~12k-token context probe.
  • The hosted auditor only reaches public https:// endpoints; audit private or internal endpoints with the self-hosted auditor.
  • On a busy card, one self-hosted run in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it.
  • When the gateway is slow, the console's availability check marks the hosted target as not running and plays its recorded audit instead (seen 2026-09-25); POST /audit/runs still ran live.
  • A pass means no evidence of a swap within the reference's measured noise, not a guarantee; small quantisation changes can stay inside the band.
  • Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the endpoint auditor API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on No GPU needed to audit. References are recorded on 1× RTX PRO 6000 (96 GB)..
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Continuous, signed checks that an OpenAI-compatible endpoint serves the model and quality it claims.

The auditor sends a fixed suite of probes at temperature 0 to any OpenAI-compatible endpoint: ten golden prompts compared with a reference run of the claimed weights (text, engine token counts and, where the endpoint exposes them, per-token logprobs), ten graded canaries, a ~12k-token context probe and a streamed throughput probe. It classifies the endpoint as pass, drift or fail, re-runs the probes before any drift or fail is signed, and signs the report with a dedicated Ed25519 key that anyone can verify. It is for teams buying open-model inference, routers and model labs who want evidence, not a provider's word, that the weights and engine behind an endpoint are the ones they pay for.

Deployment
Hosted or self-host
Regulatory
Auditing a third-party endpoint uses your own key and credits with that provider: the hosted auditor holds the key in memory for one run and never stores or logs it. Check your provider's terms before publishing results; reports are private by default. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3.5-4B-Base).
Architecture
Text description

The auditor sends golden prompts, canaries, a context probe and a throughput probe at temperature 0 to the endpoint under test. It compares the answers with a signed reference fixture recorded on our pinned stack, re-checks any drift or fail, and signs the report with its Ed25519 key. Probes to the hosted route go through our gateway, which returns a signed receipt per probe; those receipt ids are listed in the report. In self-host mode the auditor and fixtures run on your network and your key never leaves it.

Architecture

At a glance

Your provider's API key
Held in memory for one run only: never logged, stored or written into the report. The probes bill your provider (about 23-43 short requests plus the context probe).
Data retention
Signed reports are stored on the server and readable by anyone with the report id; a report of your endpoint records its URL and model name, never the key.
Hardware
CPU only: comparison data ships as signed reference fixtures. Recording new fixtures needs the model's pinned GPU stack.
Typical hosted cost
A fraction of a cent of gateway time for a short audit of our own hosted route (measured). Each run shows its own measured cost.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    audits only, no GPU

    Run the auditor against any endpoint with the signed reference fixtures we ship. You cannot record new references.

    Models
    • decosa-api auditor (decosa_api/verticals/auditor)
    • Qwen3.5-4B-Base (BF16)
    Hardware
    Any Linux or macOS machine with Python 3.11+
    Quality evidence
    • Swap caught: Qwen3.5-4B-Base served as qwen3.8-27bfail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergencesreport aud_7be2d97ae80e20abdf34, our server 2026-09-24
    • Quantisation drift caught: FP8 re-quant claimed as BF16 (4B)drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963report aud_8a7c979db05198f582de, our server 2026-09-24
    Latency
    measured: seconds per audit, depending on the endpoint and whether a re-check runs.
    Verification
    Proof: partialSigned report; probes to non-the network endpoints carry no receipts.
  • In the hosted demo

    Standard

    one 96 GB card (hosted demo)

    Adds recording your own references on the pinned Qwen3.8-27B stack, with its noise band measured idle and under load.

    Models
    • decosa-api auditor (decosa_api/verticals/auditor)
    • Qwen3.8-27B (NVFP4)
    • Qwen3.5-4B-Base (BF16)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Reference noise, Qwen3.8-27B stack (16 runs, 12 of them under load)worst greedy repeat 4/10; mean |Δ logprob| ≤ 0.055; top-5 overlap ≥ 0.826fixture qwen3.8-27b.json, our server 2026-09-30 (re-recorded after the server was restarted with image and video input on 28 Sep)
    • Hosted gateway route vs referencepass; greedy at or above the band (≥ 3/10), canaries equal to the reference. The nightly check runs this routethe nightly check (its latest result is under 'How we tested it')
    • Raw engine route vs referencepass in 10 of 10 audits on 30 Sep 2026; greedy 4 to 7 of 10 (band ≥ 3/10), top-5 overlap 0.835 to 0.893 (band ≥ 0.806)docs/evals/auditor-claims.md, 30 Sep 2026
    • Engine drift on the same weights (coding benchmark, /500)478 NVFP4 + MTP; 387 FP8 eager; 193 FP8 + MTP; 62-63 llama.cpp CUDAcoding-agent-bench README (not re-run by the auditor)
    Latency
    measured: seconds per audit on the raw route, about twice that on the receipted gateway route.
    Verification
    Proof: strongEvery probe to the hosted route has a gateway-signed receipt; the report lists their ids.
  • Best

    two 96 GB cards

    Adds a DeepSeek-V4-Flash reference so endpoints selling that model can be audited.

    Models
    • decosa-api auditor (decosa_api/verticals/auditor)
    • Qwen3.8-27B (NVFP4)
    • DeepSeek-V4-Flash (NVFP4)
    Hardware
    2x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • DeepSeek-V4-Flash reference fixturenot measured yetnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyDeepSeek-V4-Flash is not a hosted model, so its reference calls would carry no receipts.
  • Needs more compute

    Wanted: the best setup

    references for the most-used open models

    Reference runs of GLM-5.3-Flash and DeepSeek-V4.1-Flash, the two most-used open models, so endpoints that sell them can be audited against the real thing. Neither fits the reference box today. Not served yet.

    Models
    • decosa-api auditor (decosa_api/verticals/auditor)
    • GLM-5.3-Flash
    • DeepSeek-V4.1-Flash
    Hardware
    Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). Estimate.
    Quality evidence
    • substitution detection on these modelsnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Probe runner, scorer and signer (no model; runs on CPU)decosa-api auditor (decosa_api/verticals/auditor)
0 GBProof: partial
Reference model (golden outputs, logprobs, noise band)Qwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strong
Small reference model (swap and quantisation demos)Qwen3.5-4B-Base (BF16)Qwen/Qwen3.5-4B-Base on Hugging Face (opens in a new tab)
4B · 13 GBNo proof yet
Larger reference (planned)DeepSeek-V4-Flash (NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active)No proof yet
Reference for GLM-5.3-Flash endpointsGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only
Reference for DeepSeek-V4.1-Flash endpointsDeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab)
552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • verify_audit_report.pyApache-2.0

    Offline check of a signed report with only the cryptography package (decosa-api scripts/).

  • coding-agent-bench

    The owner's coding benchmark and source of the engine-drift evidence: the same Qwen3.8-27B weights scored 478, 387, 193 and 62-63 out of 500 depending on engine.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /audit/targets, POST /audit/runs (SSE), GET /audit/reports/{id}, GET /audit/signing-key, GET /audit/references, POST /audit/verify.

  • auditor-swap-4b (demo only):8201
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Serves Qwen3.5-4B-Base with FP8 on-the-fly quantisation under two names: qwen3.8-27b (the swap) and qwen3.5-4b-base (the drift). Loopback only, about 13 GB on GPU0.

Hardware

  • Any CPU, no GPU Fits

    Running audits. Reference fixtures ship with the auditor, signed.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Recording the Qwen3.8-27B reference on the pinned stack; measured on our server.

  • 2x RTX PRO 6000 Blackwell 96 GB Fits

    Needed to record a DeepSeek-V4-Flash reference. Not done yet.

Latency per lane

  • full audit, raw engine route with logprobs (23 requests, no re-check)7.7 s

    Measuredmeasured on our server 2026-09-24 (report aud_b5a3508ddb34527f490e)

  • full audit, hosted gateway route with receipts (23 requests)15.3 s

    Measuredmeasured on our server 2026-09-24 (report aud_687900df0710578a42a8)

  • full audit with re-check, 4B endpoint (43 requests)21.7 s

    Measuredmeasured on our server 2026-09-24 (report aud_7be2d97ae80e20abdf34)

  • reference decode speed, Qwen3.8-27B NVFP4 + MTP3, single stream8 ms

    Measuredmeasured on our server 2026-09-24: 129.5 tok/s median of 3 (fixture qwen3.8-27b)

Notes

  • The gateway route reports its own metered token estimates (for example 34 prompt tokens where the engine counts 26 for the same request), so the token-count fingerprint is skipped on that route; greedy match carries identity there.
  • Greedy decoding on the pinned 27B stack is not batch-invariant: under concurrent load only 5 of 10 golden texts reproduced and the mean |Δ logprob| reached 0.040. The pass band is set from that measured spread (6 runs: 3 idle, 3 with 4 concurrent streams), not from an idle run.
  • An FP8 re-quantisation of a 4B model moves logprobs by about as much as batch noise does on the 27B. It is caught through the top-5 overlap and greedy match against the 4B's own tighter band, not by the mean shift alone. Small quantisation changes on a busy 27B endpoint may stay inside its band: the auditor would say pass, and that limit is real.
  • Engine-drift evidence we did not re-run here: the same Qwen3.8-27B weights scored 478/500 on vLLM NVFP4 + MTP, 387 on vLLM FP8 eager, 193 on vLLM FP8 + MTP and 62-63 on llama.cpp CUDA (coding-agent-bench README). A second 27B engine config could not be stood up without stopping a live service.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

auditor/assemble-prompt.md118 lines
You are setting up the **Decosa endpoint auditor** on this machine. It checks whether an OpenAI-compatible endpoint really serves the model it claims (for example `qwen3.8-27b`), at the quality it claims, and signs each result with an Ed25519 key that lives only here. It needs **no GPU**: the comparison data ships as signed reference fixtures. API keys for the endpoints I audit must stay on this machine. Work step by step, show me each command before running anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/auditor.zip (16 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py auditor` (the api image carries the same bundle under /app/rehearsal/auditor/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py auditor --bundle auditor.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least one audit target is up on this server", "the audit finishes without errors", "the audit compares against the shipped reference fixture"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 1. Check the machine
- Any Linux or macOS machine with Docker, or Python 3.11+. No GPU is needed to run audits.
- If `docker` is missing, install Docker Engine from the official docs (https://docs.docker.com/engine/install/). Ask me before adding my user to the `docker` group.
- Recording **new** reference fixtures does need the GPU and the exact pinned stack of the model (see step 7). Skip that unless I ask.

## 2. Image (pinned)
- `${DECOSA_REGISTRY}/decosa-api:0.1.0` contains the auditor (`decosa_api/verticals/auditor`), its probe suite `auditor-suite-v1` and the signed fixtures `qwen3.8-27b` and `qwen3.5-4b-base`. **It is publishing soon.** Try `docker pull` first.
- If the pull fails with "not found", "denied" or "unauthorized", stop and tell me. Once the decosa-api source is published, the fallback is `docker build -f docker/api/Dockerfile -t decosa-api:local .` from its checkout. Do not substitute any other image.

## 3. Write `~/decosa-auditor/docker-compose.yml`
```yaml
name: decosa-auditor
services:
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:0.1.0
    restart: unless-stopped
    ports: ["127.0.0.1:8445:8445"]          # localhost only
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct              # the auditor itself calls no model of ours
      DECOSA_AUDITOR_KEY: /data/keys/auditor_ed25519.key   # inside the named volume
      DECOSA_AUDITOR_CREATE_KEY: "1"        # first start creates the key, mode 0600
      DECOSA_AUDIT_ALLOW_PRIVATE: "1"       # self-host mode: audit endpoints on my own network too
      DECOSA_AUDIT_TARGETS_JSON: /config/targets.json
    volumes:
      - auditor-data:/data                  # audits.sqlite (signed reports) and keys/ (the signing key)
      - ./targets.json:/config/targets.json:ro
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request;urllib.request.urlopen('http://127.0.0.1:8445/audit/signing-key',timeout=3)"]
      interval: 15s
      timeout: 5s
      retries: 10
volumes:
  auditor-data:
```
- The signing key lives in the named volume, not in a host folder: the image runs as an unprivileged user (uid 10001)
  and could not write a `./keys` folder, so the key was never created and `/audit/signing-key` answered 503. After the
  first start, back it up with `docker compose cp api:/data/keys ./auditor-keys-backup` (mode 0700, never commit it).

## 4. Write `~/decosa-auditor/targets.json`: the endpoints to audit
Ask me for each endpoint: its base URL (ending in `/v1`), the model name it serves, and the model it claims to be. Example:
```json
[
  {"id": "prod-a", "kind": "openai", "base_url": "http://10.0.0.12:8000/v1", "model": "qwen3.8-27b",
   "claimed_model": "qwen3.8-27b", "label": "Inference box A"}
]
```
Do not put API keys in this file. Keys are passed per run (step 6) and are held in memory only.

## 5. Start and check health
- `cd ~/decosa-auditor && docker compose pull && docker compose up -d`
- Poll `curl -fsS http://127.0.0.1:8445/audit/signing-key` until it returns `{"alg":"ed25519","pubkey":"…","key_id":"aud-…"}`. Show me the `pubkey`: it is what anyone pins to verify my reports.
- `curl -fsS http://127.0.0.1:8445/audit/references` should list `qwen3.8-27b` and `qwen3.5-4b-base`, each with `fixture_sha256` and `fixture_sig`.
- `curl -fsS http://127.0.0.1:8445/audit/targets` shows each target with `"available": true|false`.

## 6. Smoke test: run one audit
```bash
TOKEN=$(curl -fsS -X POST http://127.0.0.1:8445/demo/session -H 'Content-Type: application/json' \
  -d '{"vertical":"auditor"}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["token"])')
curl -N -X POST http://127.0.0.1:8445/audit/runs -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
  -d '{"target":"prod-a"}'
```
- For an endpoint that needs a key, send `{"base_url": "...", "model": "...", "claimed_model": "...", "api_key": "$KEY", "context_tokens": 4000}` instead, reading `$KEY` from an environment variable; never echo it.
- Expect SSE events: `ready`, then `lane` events for `identity`, `quality` and `performance` (each carries `data.checks`, the lane's full current state), a `verdict` lane, and `done` with `summary.report`.
- A run is about 23 requests (43 when a re-check runs) and about 1.5k generated tokens plus the context probe. On a paid provider that is the cost of one audit.
- Save the report: `curl -fsS http://127.0.0.1:8445/audit/reports/<id> > report.json` and verify it with the snippet below.

## 7. Verify a report (anyone can do this)
```python
import json
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
report = json.load(open("report.json")); pub = "<the pubkey from step 5>"
assert report["signer"]["pubkey"] == pub
body = {k: v for k, v in report.items() if k != "signature"}
msg = b"decosa.audit.report.v1\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(report["signature"]), msg)
print(report["verdict"])
```

## 8. How to read a verdict
- **pass**: within the reference stack's own measured spread (idle and under load) on identity and quality.
- **drift**: same family, different numerics or quality: quantisation, engine or kernels. **fail**: does not behave like the claimed weights. Both are signed only after a second pass agrees; otherwise the verdict is **inconclusive**.
- Endpoints that hide logprobs are judged on greedy match and token counts alone, which is weaker evidence. The report says so.
- Small quantisation changes on a busy endpoint can stay inside the noise band. A pass is "no evidence of a swap", not a guarantee.

## 9. Point other tools at it
- Schedule audits with cron or a systemd timer calling step 6, and alert on `verdict.status` other than `pass`.
- Only use the fixtures we ship, or ones you record yourself on the model's exact pinned stack with `scripts/auditor_record_reference.py` (GPU needed; it measures the noise band idle and under concurrent load).

When done, report: the auditor's `pubkey` and `key_id`, the references listed, each target's availability, and the verdict of the smoke-test audit.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Continuous, signed checks that an OpenAI-compatible endpoint serves the model and quality it claims.
Who it's for
Teams in software and ai ops and compliance and trust.
Where it runs
Hosted or self-host
Key numbers
  • fail, re-check agreed; greedy 0/10, top-5 overlap 0.551, 7 hard divergences Swap caught: Qwen3.5-4B-Base served as qwen3.8-27b (synthetic)
  • drift, re-check agreed; greedy 5/10, top-5 overlap 0.928 vs band ≥ 0.963 Quantisation drift caught: FP8 re-quant claimed as BF16 (4B) (synthetic)
  • worst greedy repeat 5/10; mean |Δ logprob| ≤ 0.040; top-5 overlap ≥ 0.872 Reference noise, Qwen3.8-27B stack (6 runs) (synthetic, n = 6)
  • 39.8 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Reference runs of Qwen3.8-27B · Qwen3.5-4B
Where
Hosted or self-host
Checks
Signed report, receipted probes
Output
Signed record or verdict
Data
No sensitive data
Hardware
CPU, no GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

How does the Decosa endpoint auditor detect a model swap?

The Decosa endpoint auditor sends a fixed suite of probes at temperature 0 to any OpenAI-compatible endpoint: ten golden prompts compared with a reference run of the claimed weights (text, token counts and, where exposed, per-token logprobs), ten graded canaries, a ~12k-token context probe and a streamed throughput probe. It classifies the endpoint as pass, drift or fail and re-runs the probes before any drift or fail is signed. A pass means no evidence of a swap within the reference's measured noise, not a guarantee.

Can the auditor tell if a provider serves a quantized model?

In a planted test, the Decosa endpoint auditor flagged an FP8 re-quantised 4B model claimed as BF16 as drift, and the re-check agreed: greedy 5/10, top-5 overlap 0.928 against a band of at least 0.963. A smaller model served under a larger model's name failed with greedy 0/10 and top-5 overlap 0.551. Only these two swaps were tested, both set up by the builder, and small quantisation changes can stay inside the noise band.

What happens to my provider API key during an audit?

The hosted Decosa endpoint auditor holds your provider's API key in memory for one run only; it is never logged, stored or written into the report. The probes bill your provider, about 23-43 short requests plus the context probe. Signed reports are stored on the server and readable by anyone with the report id; a report records the endpoint URL and model name, never the key. Check your provider's terms before publishing results.

Does the endpoint auditor need a GPU?

No. The Decosa endpoint auditor runs on CPU because comparison data ships as signed reference fixtures; the Lite tier runs on any Linux or macOS machine with Python 3.11+. Recording new reference fixtures needs the model's pinned GPU stack, for example one 96 GB RTX PRO 6000 for Qwen3.8-27B. A DeepSeek-V4-Flash reference is planned but not measured yet. The hosted auditor only reaches public https endpoints; audit private ones self-hosted.

How reliable is a single audit run?

A single run can be noisy. On a busy card, one self-hosted Decosa endpoint audit in three came out inconclusive: the first pass flagged drift (top-5 overlap 0.849 against a band of 0.852) and the re-check did not reproduce it. That is why drift and fail are re-run before they are signed. A short hosted audit of our own route cost about $0.002 of gateway time (22 probes) on 2026-09-25.

How can someone else check an audit report?

Each Decosa endpoint audit report is signed with a dedicated Ed25519 key that anyone can verify. In a self-hosted test on 2026-09-25, a full audit of a local Qwen3.8-27B vLLM returned pass in 29 s with 23 probes, the report verified with the documented Python snippet, and verification failed after the report was edited. The signature proves the report is unchanged, not that the endpoint will behave the same tomorrow.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Endpoint auditor

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.