# Decosa demo API — contract v0 (2026-09-23)

The public site talks to one service, `decosa-api`, at `https://api.decosa.ai`. The frontend reads the base URL from `NEXT_PUBLIC_DECOSA_API`.

## Verticals (ids)
`clinical` · `sales` · `code` · `studio` · `field` · `translate`

## Auth: demo sessions
- `POST /demo/session` with body `{ "vertical": "<id>" }` returns `{ "token": "<opaque>", "expires_at": <unix>, "budget": { "seconds_audio": 300, "llm_tokens": 20000 } }`.
- Rate limits: 20 sessions per network per hour, 60 per hour in total (`GET /healthz` reports both under `demo_sessions`). Global concurrency cap: live audio sessions ≤ 4 (configurable). When over the cap, return 429 with `Retry-After`.
- Every other call sends `Authorization: Bearer <token>`. WebSockets take `?token=` instead.
- No accounts and no PII are kept. Session transcripts live in memory only and are deleted when the session ends. The clinical demo shows a banner: "Demo only. Do not enter real patient information."

## Live audio verticals (clinical, sales, field, translate)
`WS /ws/live?vertical=<id>&token=<t>[&lang=<src>&target=<dst>]` (lang and target apply to `translate` only)

**Client → server:**
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms per frame;
- text frame `{"type":"stop"}`.

**Server → client:** JSON text frames.
- `{"type":"ready","vertical":"clinical","models":{"asr":"voxtral-mini-4b-realtime","llm":"qwen3.8-27b"}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true}` (partial updates have `final:false`)
- `{"type":"lane","lane":"<lane_id>","title":"<Title>","body":"<markdown or text>","data":{...optional structured...},"latency_ms":850}`
- `{"type":"receipt","id":"<completion id>","model":"qwen3.8-27b","gateway_sig":"<hex>","provider":"<provider id>"}`
- `{"type":"budget","seconds_audio_left":240,"llm_tokens_left":15000}`
- `{"type":"error","message":"..."}`
- `{"type":"done","summary":{...final artifact...}}`

**Lane ids per vertical:**
- clinical: `note` (SOAP), `codes` (ICD-10-CM + RxNorm, validated), `priorauth`, `visit_level`; final artifact = note plus codes.
- sales: `objection`, `next_question`, `crm` (JSON fields), `followup` (email, sent on stop).
- field: `checklist`, `issues` (list with severity), `measurements`, `report` (on stop: structured report JSON plus markdown).
- translate: `translation` (streamed per sentence, target language), `glossary`, `summary` (on stop).

**Text fallback, for demos without a mic:** `POST /demo/replay` with `{ "vertical": "<id>", "script_id": "<id>" }` streams the same event types over SSE, using a canned audio script run through the REAL pipeline, so the output is live.

## Code assistant
`POST /v1/chat/completions` is OpenAI-compatible and streams (`stream:true`, SSE). Model `qwen3.8-27b`, bearer = the demo token. Each response carries header `x-decosa-receipt: <completion id>`. The final SSE chunk includes `"receipt":{...}` in the same shape as the WS receipt event.

## Studio
- `GET /studio/gallery` returns `[{ "id", "kind": "music"|"image"|"video", "title", "prompt", "model", "url", "seed", "duration_s"? }]`. Pre-rendered outputs are static files served under `/studio/media/...`.
- `POST /studio/jobs` with body `{ "kind", "prompt", "params":{} }` returns `{ "job_id", "status":"queued", "position": N }`.
- `GET /studio/jobs/{id}` returns `{ "status": "queued"|"running"|"done"|"failed", "url"?: ... }`.

## Watch (recorded sessions)
- `GET /demo/recordings` returns `[{ "id", "vertical", "title", "duration_s", "events_url", "video_url"? }]`.
- `GET <events_url>` returns a JSON array of the same event objects as the WS stream, each with an `at_ms` offset. The frontend replays them with a player (scrubber, play/pause, 1x/2x).

## Receipts
`GET /receipts/{id}` returns `{ "id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id","pubkey","sig"}, "gateway": {"pubkey","sig"}, "proof": {"format","verified":true}, "checks": [{"name","ok","detail"}] }`. Read-only and public. It proxies our gateway's receipt data and never exposes operator routes.

## Health
`GET /healthz` returns `{ "ok": true, "asr": true|false, "llm": true|false, "live_sessions": n, "queue": n }`

## For developers ("build with it")
Each vertical page has a copy-paste prompt for a coding agent. It describes the API above (base URL, auth, events, example code). The clinical and field pages also cover self-hosting (`decosa` CLI / docker compose) and say the hosted API must not receive PHI.

## Changes (decosa-api v0.1.0, 2026-09-23)
These are clarifications and additions made while implementing the contract. Nothing listed above was removed.

**WebSocket close codes** (an `error` event is always sent first): `4401` bad or expired token; `4429` busy (live cap reached, retry in 30 s); `4400` bad parameters (vertical, `lang`/`target`); `4402` session budget used up; `4403` token is for another vertical, or `Origin` not allowed; `4409` this token already has a live session; `1011` ASR unavailable; `1000` normal end after `done`.

**Transcript events:** `final:false` carries the cumulative text of the sentence in progress. `final:true` carries the finished sentence and replaces it. Both carry `i` (sentence index) and `t` (seconds since the session started, at the sentence's first word).

**Lane events:** the latest event for a lane id is that lane's full current state, so replace, don't append. The exception is `translation`: one event per source sentence, ordered by `data.i` (`data` = `{i, t, source, translation, target}`), and they can arrive out of order under load. Every event from the server also carries `at_ms`, milliseconds since the session started.

**Receipt event** adds `"status": "signed" | "unverified"`. `unverified` appears only when the server runs with `DECOSA_LLM_ROUTE=direct`, which means no gateway receipt exists and none is ever invented.

**`GET /receipts/{id}`** adds `status`, `upstream_model`, `usage`, `cost` and `receipt`. `receipt` is the raw gateway-signed receipt: hashes and signatures only, no text, so anyone can re-verify it. `weights_root` is the receipt's `model_manifest_hash`. For first-party runtimes, `provider.pubkey` and `provider.sig` are `null`, because no separate provider key exists. `checks` = commitment hash, gateway Ed25519 signature, gateway key pinned to the key the gateway publishes, sidecar metering proofs, and a provider signature when one is present.

**`ready` event** adds `receipts: "signed" | "unverified"`. Clinical adds `banner` (the demo-only text) and `locale: "us"`. Translate adds `lang` and `target`. `lang` defaults to `auto` and `target` to `en`. Both must be language codes like `es` or `pt-BR`.

**Budgets:** `llm_tokens` counts generated (completion) tokens. Prompt size is bounded separately: the code route accepts at most 48,000 characters per request (413 above that), and the live lanes build their own prompts. Live lanes keep part of the budget back for the stop artifact. When only that reserve is left, live lanes pause and an `error` event says so. The session ends when the audio budget or the token budget is used up.

**HTTP status codes:** `POST /demo/session` returns 429 + `Retry-After: 30` for a live vertical when the live cap is full, and 429 + `Retry-After` (seconds until the oldest of the address's sessions leaves the one-hour window) at the per-address limit. `/demo/replay` returns 401 for a missing or invalid token, 403 when the token is for another vertical, 404 for an unknown script, 402 when the budget is used up, 409 when the token is already live, and 429 + `Retry-After: 30` when the cap is full. `/v1/chat/completions` returns 402 when the budget is used up, 413 for an oversized prompt and 502 when the upstream fails.

**Replay:** the body also accepts `speed` (1.0–2.0, default 1.0). The audio is paced in real time through the same ASR and LLM pipeline. Script ids:
- `clinical-back-pain`, `clinical-diabetes-followup`
- `sales-discovery-logistics`, `sales-renewal-security`
- `field-roof-inspection`, `field-electrical-panel`
- `translate-es-en-repair` (es→en), `translate-en-es-tour` (en→es)

New `GET /demo/scripts` (no token) lists them as `[{id, vertical, title, lang, target, duration_s, has_audio}]`. Recording ids equal script ids. `events_url` is a path relative to the API base.

**Code assistant:** also `GET /v1/models`. Any valid demo token works. With `stream:true` the server gets the completion from the gateway's metered route, then re-emits it as SSE chunks, because the gateway's streaming route settles a receipt but does not return one. So time-to-first-token equals the full generation time. Tools/function calling is not supported on the hosted route. `max_tokens` is capped at 2048 and by the remaining budget.

**Studio:** `POST /studio/jobs` needs a `studio` token, and each token can submit at most 3 jobs. `GET /studio/jobs/{id}` (no token) adds `position` and, when no render worker is running, a `note`. The gallery lists only files that exist under `/studio/media/`. It is empty until real renders are added.

**No token needed for:** `/healthz`, `/demo/scripts`, `/demo/recordings`, `/demo/recordings/{id}/events`, `/studio/gallery`, `/studio/media/...`, `/studio/jobs/{id}`, `/receipts/{id}`. `/healthz` also returns `llm_route` and `version`.

**CORS:** allowed origins are `^https://decosa(-[a-z0-9-]+)?\.vercel\.app$` plus `localhost`/`127.0.0.1` on any port. They are configurable with `DECOSA_CORS_ORIGINS` and `DECOSA_CORS_ORIGIN_REGEX`. The exposed headers are `x-decosa-receipt` and `Retry-After`. WebSockets check `Origin` against the same list. A request with no `Origin`, i.e. not from a browser, is allowed.

**After-visit pass 2 (added 2026-09-23).** When a clinical session ends with `{"type":"stop"}` (or a replay finishes), the server runs the scribe-bench two-pass pipeline over the whole session audio, which it kept in memory:
1. MOSS-Transcribe-Diarize 0.9B (Apache-2.0) produces a speaker-attributed transcript.
2. An LLM role map turns S01/S02 into Doctor/Patient.
3. Qwen3.8-27B writes the note with `[[n]]` span citations.
4. A claim-level LLM verifier checks each claim.

It emits three new `lane` ids before `done`:
- `final_transcript` (title "Transcript (diarized)"). `data` = `{lines: [{n, speaker, role, start, end, text}], roles: {S01: "Doctor", …}, model, audio_s, diarize_s}`. `n` is the line number that citations refer to.
- `final_note` (title "Visit note (cited)"). `body` is the note text with sections CHIEF COMPLAINT, HISTORY OF PRESENT ILLNESS, PHYSICAL EXAM, RESULTS, ASSESSMENT AND PLAN. Every sentence ends in `[[n]]` or `[[n,m]]`. `data` = `{note, claims: [{i, section, claim, cites}]}`.
- `verifier` (title "Claim check"). `data` = `{counts: {SUPPORTED, PARTIAL, UNSUPPORTED, ERROR}, flagged, claims: [{i, section, claim, cites, evidence, label, leak, reason}]}`. A claim is flagged when its label is PARTIAL or UNSUPPORTED, or when `leak` is true (health information about someone other than the patient). `ERROR` means the judge's answer couldn't be parsed; the claim is never counted as supported.

The clinical `done.summary` adds `final_transcript`, `final_note`, `verifier` and `pass2_timing` (`{diarize_s, roles_s, note_s, verifier_s, total_s}`). With pass 2, `summary.note` is the cited note. If pass 2 fails, an `error` event says so, the note falls back to the single-pass SOAP note written from the live transcript, and `summary.pass2` gives the reason.

The sales vertical runs steps 1–2 only (roles: Seller/Buyer). It emits `final_transcript`, writes the follow-up email from the attributed transcript, and adds `final_transcript` and `pass2_timing` to `done.summary`.

Every pass-2 LLM call goes through the gateway with its own `receipt` event and counts against the session budget. The clinical pack reserves 5,000 generated tokens for it. `/healthz` adds `diarize: true|false`. The diarizer is a separate local service (`services/diarize`, 127.0.0.1:8092, GPU0). Audio is never written to disk.

**Developer API keys (added 2026-09-23).** A developer key (`dk_…`) works alongside demo session tokens. Every use case offers "self-host" or "get an API key".

Admin routes require `Authorization: Bearer <DECOSA_ADMIN_SECRET>`. They are called server-to-server by the site and are never exposed to browsers. Without the secret they return 401; if no secret is configured, 503.
- `POST /v1/keys`
  - Body: `{label?, verticals?: [ids] | "all", daily_llm_tokens?, daily_audio_seconds?, plan?}`.
  - Returns 201 `{key, key_id, created, label, plan, limits: {verticals, daily_llm_tokens, daily_audio_seconds, requests_per_minute, concurrent_live_sessions}}` with `Cache-Control: no-store`.
  - The `key` appears only in this response. The server stores only its SHA-256.
  - Defaults (free tier): all verticals, 200,000 generated LLM tokens a day, 1,800 audio seconds a day, 60 requests a minute, 2 concurrent live sessions. Caps: 5M tokens and 8 h of audio a day. `plan` defaults to `"free"` and is reserved for credit billing later.
- `GET /v1/keys/{key_id}` → `{key_id, label, plan, created, revoked, revoked_at, limits, today: {day, llm_tokens, audio_seconds, requests, llm_tokens_left, audio_seconds_left}, history: [last 30 days]}`. Days are UTC.
- `DELETE /v1/keys/{key_id}` → `{key_id, revoked: true}`. Revocation is immediate, and later uses get 401.

Using a key: a `dk_` key goes anywhere a demo token does. REST takes `Authorization: Bearer dk_…`; WebSockets take `?token=dk_…`. The only exception is `POST /demo/session`, which isn't needed with a key.
- A key restricted to some verticals gets 403 elsewhere, and WS close `4403`.
- Budgets are daily, not per session. An exhausted budget returns 402 on REST and close `4402` on WS. The error message ends "for today".
- Over the request rate limit: 429 with `Retry-After`, and WS close `4429`. Over the per-key live cap: 429 on `/demo/replay`, and WS close `4429`. Replays and live WebSockets both count toward the cap.
- The global live-session cap still applies to keys.
- `budget` events on a key session add `key_id` and `period: "day"`.
- Every LLM call still goes through the gateway and gets a signed receipt. The `receipt` event and `GET /receipts/{id}` add `key_id`, which is `null` for demo tokens.

Storage: keys and daily usage live in SQLite at `<DECOSA_DATA_DIR>/keys.sqlite` (0600).
**Studio (updated 2026-09-23).**

Gallery:
- Items can carry `notice`, e.g. "Powered by MiniMax-Music3". It must be shown wherever the item is shown.
- Every gallery item carries `report_url` (`/studio/report?item=<id>`).

Jobs:
- `POST /studio/jobs` (`params`: `seed`, `seconds`, `lyrics` for music, `aspect` = `square` | `landscape` | `portrait` for images, and `engine: "acestep"` to use ACE-Step for music).
  - Before queueing, the prompt and the lyrics go through the prompt policy. A rejection returns 422 `{error, category, policy_url: "/studio/policy"}`.
  - `category` is one of `real_artist`, `real_song`, `voice_imitation`, `copyrighted_character`, `infringing` or `unavailable`. The LLM layer fails closed, so `unavailable` means "try again".
- The job response and `GET /studio/jobs/{id}` add `report_url` (`/studio/report?job=<id>`). They also add `notice` on finished jobs, and `paused: true` plus a `note` while a live demo session holds the GPU.
- Routing: video → Wan2.1-T2V-14B (832×480, up to 5 s at 16 fps); music → MiniMax-Music3 (up to 60 s), or ACE-Step 1.5 when `engine: "acestep"`; image → Qwen-Image-2512. LTX and MiniMax-H3 are never used for hosted jobs.
- Jobs run one at a time and don't start while a live session is running. A running render is paused during live sessions.

New endpoints:
- `GET /studio/policy` returns the rules as text and the report categories.
- `GET /studio/report?item=|job=` describes the form.
- `POST /studio/report` takes `{item | job, category, details?, contact?}`.
  - `category` ∈ `copyright`, `voice_imitation`, `real_person`, `harmful`, `privacy`, `other`.
  - Returns 201 `{report_id, received}`.
  - Errors: 400 for bad input, 404 for an unknown item or job, 429 above 5 reports per hour per address.
  - Reports are stored server-side (SQLite, 0600). Logs carry only ids and hashes.

**Field pack: grounded report, safety sweep, claim check (added 2026-09-23).** The four v0.1 field lane ids (`checklist`, `issues`, `measurements`, `report`) are unchanged, and every v0.1 field is still there. Additions:
- The lane model sees the transcript as numbered lines `n [m:ss] text` and cites them. Each issue, measurement and passed check carries `quote` (verbatim transcript text; the model's quote is kept only when it matches the transcript, otherwise the cited line's own text is used), `lines` (the cited line numbers) and `t` (seconds since the session started, at the first cited line). `issues` items keep `evidence`, now equal to `quote`. Measurements add `spec`: the limit, expected range or threshold as the inspector stated it, or empty.
- New tick lane `passed_checks` (title "Passed checks"): `data` = `{items: [{item, quote, lines, t}]}`, the things checked and found fine.
- On stop, three lanes arrive before `done`, in this order:
  - `safety_sweep` (title "Safety sweep"). The transcript is checked against a checklist of common safety items for the inspection type (electrical: GFCI/AFCI, grounding and bonding, panel clearance, double taps and others; roof: flashing, sealant, penetrations, gutters and others; also HVAC, fire, plumbing, structural, general). Anything the inspector mentioned that the report left out is added to the report. An addition must cite a real transcript line, and anything that doesn't is rejected, so the sweep never adds an item nobody mentioned. `data` = `{categories, checklist, added_issues, added_passed, rejected: [{issue, why}]}`. Added issues carry `source: "safety_sweep"`.
  - `report_check` (title "Claim check"). This is the pass-2 claim verifier (`pass2.verify`) with an inspection-auditor prompt. It judges every cited report item and every summary sentence against its transcript lines (one call each; a verdict that can't be read is asked once more). `UNSUPPORTED` items, and items that cite no transcript line, are removed from the report and listed in `removed`. `PARTIAL` items are rewritten to the judge's correction when it changes something without making the statement longer; they are listed in `corrected` and marked `verified: "corrected"`. Otherwise they are kept with `verified: "partial"` and a `check_note`. A cited item whose verdict still can't be read stays, marked `verified: "unchecked"`. Summary sentences that are unsupported or unjudged are left out. `data` = `{counts, removed: [{section, claim, label, reason, quote, lines}], corrected: [{section, was, now, reason}], claims: [{i, section, claim, cites, label, reason, corrected?}]}`.
  - `report`, as before. `data.report` adds `inspection_type`, `site_details: [{fact, quote, lines, t}]`, `passed_checks: [{item, quote, lines, t}]` and `code_references: [{reference, applies_to, quote, lines, t}]` (only codes or requirements the inspector stated). Every issue, measurement, site detail, passed check and code reference has `quote`, `lines`, `t` and `verified` (`supported` | `corrected` | `partial` | `unchecked`). `summary` is plain text; its sentences were each checked. The markdown adds a Spec column, Passed checks, Code references, the quote and time for each issue, and a "Removed by the claim check" list for review.
- `done.summary` for field adds `safety_sweep`, `report_check`, `transcript_lines: [{n, t, text}]` (the numbering that `lines` refers to) and `timing: {report_s, sweep_s, check_s}`.
- Every sweep and claim-check call goes through the session like any other LLM call, with its own `receipt` event and a charge to the budget. The field pack now reserves 5,500 generated tokens for the stop lanes (it was 1,400).

**Attested receipts, session hash chains and the record vertical (added 2026-09-23).**

Shared primitives (any vertical can use them):

*1. This server's own key.* decosa-api now has an Ed25519 key (`attest.py`), generated on first start at `<DECOSA_DATA_DIR>/attest/ed25519.pem` (0600, or `DECOSA_SIGNING_KEY`). `DECOSA_LOCAL_SIGNING=off` turns it off.
- `GET /attest/signing-key` (no token) → `{scheme:"ed25519", pubkey, key_id, name, created, signs:[domains], note, llm_route, receipts}`. 404 when signing is off.
- `GET /attest/models` (no token) → `{<served name>: {repo, revision, root, scheme:"hf-files-v1", files, bytes, source, manifest:[{path, sha256, size}]}}`. `root` = sha256 of the canonical JSON `{scheme, repo, revision, files}` (`scripts/model_root.py`; anyone can recompute it from the Hub's per-file SHA-256). Today: Voxtral Mini 4B Realtime, MOSS-Transcribe-Diarize, Qwen3.8-27B NVFP4 (self-host revision). On the gateway route the LLM's root is instead the gateway receipt's `model_manifest_hash` (the network's manifest).
- Signed statements are canonical JSON: sorted keys, no whitespace, UTF-8 (not ASCII-escaped), **no floats** (times are integer ms). The bytes signed are `<domain> + "\n" + canonical(statement without "sig")`. Domains: `decosa.asr-receipt.v1`, `decosa.llm-receipt.v1`, `decosa.chain-checkpoint.v1`, `decosa.record.v1`.
- **Honest scope:** these are attestations by the operator of the box that signed them. They prove the hashes weren't changed after signing and which key signed. They do not prove the model ran, that speech was recognised correctly, or that the operator is honest. Gateway receipts add a second signer and metering proofs; these do not.

*2. ASR receipts.* A pack with `asr_receipts = True` (today: `record`) signs one receipt per finished live caption, and the record pack signs one per pass-2 line:
`{v:"decosa.asr-receipt.v1", id:"asr-…", session, seg, stage:"live"|"final", model, model_root, audio_sha256, audio_bytes, audio_format:"pcm_s16le_16k_mono", audio_start_ms, audio_end_ms, text_sha256, created_ms, signer, speaker?, sig}`.
- `stage:"live"`: the audio is everything the server received between the previous caption and this one (a streaming recogniser gives no word alignment). The live spans tile the session audio.
- `stage:"final"`: the exact span the diarizer timed, `[start, end)` of the session PCM.
- They arrive as `receipt` events with `kind:"asr"`, `status:"attested"`, `signer`, `sig`, and are served at `GET /receipts/{id}` with `kind`, `status:"attested"`, `request_hash` = audio hash, `output_hash` = text hash, `weights_root` = model root, `signer:{pubkey,sig,key_url}`, `audio`, `note`, `checks:[signature, signer_pinned]`, and the raw `receipt`.

*3. Direct-route (self-host) LLM receipts.* On `DECOSA_LLM_ROUTE=direct` every LLM call now gets `{v:"decosa.llm-receipt.v1", id, model, upstream_model, model_root, request_sha256, response_sha256, prompt_tokens, completion_tokens, created_ms, signer, sig}`, `response_sha256` = sha256 of the output text, `request_sha256` = sha256 of canonical `{messages, params}`. The receipt event's `status` is `"attested"` (was `"unverified"`); `unverified` remains only with local signing off. The `ready` event's `receipts` is `"signed"` (gateway) | `"attested"` (direct + key) | `"unverified"`. On the gateway route nothing changes: receipts stay gateway-signed. Checked on our server 2026-09-23: a gateway receipt's `response_hash` equals sha256 of the completion text, so a stored output can be matched to its gateway receipt.

*4. Session hash chains* (`hashchain.py`). A pack with `chain = True` keeps an append-only chain per session. Entry: `{seq, kind, at_ms, prev, …fields, text_sha256?, hash, text?}`, `hash = sha256("decosa.chain-entry.v1\n" + canonical(entry without hash and text))`, `prev` of entry 0 = 64 zeros. Text sits beside its hash, so a record can be shared without the words. Every 25 entries and at the end the chain signs a checkpoint `{v:"decosa.chain-checkpoint.v1", session, seq, head, root, signed_ms, signer, sig}`. `root` is a Merkle root over entry hashes (leaf = sha256(0x00‖h), node = sha256(0x01‖l‖r), an odd node carried up); `SessionChain.proof(seq)` gives an inclusion proof. `seal()` returns the record:
`{format:"decosa.record.v1", statement:{v, session, vertical, title, started_ms, sealed_ms, count, head, root, models, signer, signer_name, signer_key_id, …}, sig, entries, checkpoints, note}`.
The engine chains automatically: `audio` (sha256 per 5 s of received PCM), `caption` (with its ASR receipt), `llm` (every model call: `lane`, `receipt_id`, `status`, `model_root`, `output_sha256`, plus the full `gateway_receipt` or attested `receipt`), and `lane` (every lane body except `chain` and `record`). Packs add their own kinds.

*5. Verification.* `POST /record/verify` (no token; body = the record, or `{record}`; up to `DECOSA_RECORD_MAX_BYTES`, 8 MiB; 60 per minute per address) → `{ok, checks:[{name, ok, detail}], bad:[{seq, what, problems}], first_bad, summary, session, signer, count, issued_here}`. Checks: `entries` (each text hash, link, entry hash, embedded ASR/LLM receipt signature, gateway receipt commitment and countersignature against the pinned gateway key, and that a receipt covers the chained text/output), `count`, `head`, `root`, `signature`, `checkpoints`, `signer_pinned` (signed by this server's key; not required for `ok`). Changing one word yields `first_bad` = that entry and a summary like "Verification failed at entry 57, transcript line 12 (Clerk): text was changed". The site verifies the same records in the browser (WebCrypto Ed25519) with `src/lib/record-verify.ts`.

**Record vertical (`record`, live).** `WS /ws/live?vertical=record` and `/demo/replay`. Scripts: `record-council-meeting` (91.7 s), `record-investigation-interview` (58.8 s), synthetic, TTS.
- `ready` adds `asr_receipts:"attested"|"none"`, `signer:{pubkey,key_id,name}` and `record:{format, verify:"/record/verify", signing_key:"/attest/signing-key"}`.
- Live: `transcript` events are the captions; a `receipt` event with `kind:"asr"` follows each final one. Lanes: `actions` (tick; `data.items: [{type: motion|vote|decision|action|allegation|response|exhibit, text, who, status}]`), `chain` (`data` = `{session, count, head, kinds, checkpoints, last_checkpoint, signer}`; replace on each update).
- On stop: `final_transcript` (as clinical, roles are free-form such as "Chair", "Clerk", "Interviewer"; `data.receipts` = number of pass-2 ASR receipts; falls back to the live captions with `source:"live captions"` when the diarizer is down), `summary` (sections SUMMARY, DECISIONS AND VOTES, ACTION ITEMS, OPEN QUESTIONS; every sentence ends in `[[n]]`; `data` = `{text, claims}`), `verifier` (same shape as clinical, with a meeting/interview judge prompt), then `record` (`data` = the sealed record). `done.summary` = `{actions, final_transcript, summary, verifier, record, record_check:{ok, summary}, timing:{diarize_s, roles_s, summary_s, verifier_s, total_s}, stats}`.
- Record entry kinds: `session` (models and roots, signer), `audio`, `caption`, `llm`, `lane`, `line` (`n, speaker, role, start_ms, end_ms, source:"pass2"|"live", receipt?`), `summary` (the raw summary text; its sha256 equals the summary call's `output_sha256`, and on the gateway route the gateway receipt's `response_hash`), `sentence` (`i, section, cites, label`), `end` (`audio_ms, captions, lines, asr_receipts`).
- Budget: the pack reserves 7,000 generated tokens for the stop work. Measured on our server 2026-09-23 (gateway route, one session at a time): 91.7 s council meeting, `done` 32–39 s after the audio ended, 131 entries, record 162 KB; 58.8 s interview, `done` 11 s after, 76 entries.
- Not a certified court record. The verifier checks the summary against the transcript, not the transcript against the audio.

**Provenance kit: render receipts, C2PA credentials, file check and consent.**
Code: `decosa_api/verticals/provenance/`. Needs the `provenance` extra (`c2pa-python`, `pillow`); the watermark needs the `watermark` extra (`trustmark`, pulls torch). Without them the API still starts: stamping writes receipts only, the check reports credentials as unreadable. `DECOSA_PROVENANCE=0` turns the kit off. Signing material for C2PA: `scripts/provenance_devcert.py` writes a self-made CA and an ES256 leaf to `DECOSA_PROVENANCE_DIR` (default `<API config dir>/provenance`). Receipts and consent records are signed with the attest key above.

*Render receipt* (signed statement, domain `decosa.render-receipt.v1`, same canonical rules, floats written as decimal strings):
`{v, type:"decosa.render-receipt", id:"rr_<12 hex>", kind, digital_source_type:".../trainedAlgorithmicMedia", source:{type:"gallery"|"studio-job", id}, model:{key, model, license, revision, engine, weights_digest, weights_digest_scheme, weights_files}, render:{seed, params, params_sha256, prompt_sha256, inputs:{negative_prompt_sha256?, lyrics_sha256?}}, output:{sha256, bytes, mime}, delivered:{watermark:{scheme, payload, note}|null, pixel_sha256|null, sha256_before_credential}, consent:{id, record_sha256, likeness, expires_at}|null, determinism, notice, rendered_at, stamped_at, gateway:null, signer, sig}`.
- `id` = `rr_` + first 12 hex of sha256(`kind|model_key|output.sha256`), so it fits the 61-bit watermark payload and is stable for the same raw output.
- `weights_digest` = sha256 over sorted `path\tsha256\n` lines of every weight file the engine loads (wiki page 21 `files_sha256_digest`); computed by `scripts/provenance_model_manifests.py`. No MWR root exists for these diffusion models.
- No prompt text, lyrics or consent clip is stored or signed; only hashes.

*C2PA manifest* embedded in every stamped file (PNG, JPEG, WebP, MP3, WAV, FLAC, MP4, MOV): `c2pa.actions.v2` with one `c2pa.created` action (`digitalSourceType` trainedAlgorithmicMedia, `softwareAgent` = model id and revision), `ai.decosa.render-receipt` (the signed receipt) and, when consent was used, `ai.decosa.consent-link`. Signed ES256. **With the dev CA, public validators report state `Valid` with failure code `signingCredential.untrusted`** (checked with c2patool 0.27.22 on an image, a song and a video, 2026-09-23). A trust-list certificate is an owner decision.

*Studio integration.* Every finished job is stamped in place: `GET /studio/jobs/{id}` adds `receipt_id`, `receipt_url`, `credential:"c2pa"`, or `credential_error` when stamping failed (the render still stands). Gallery rows with a receipt add `receipt_id`, `receipt_url`, `credential_url`, and `url` points at the credentialed copy (`original_url` keeps the unstamped file). The 15 gallery items were backfilled from their recorded seeds and settings (`scripts/provenance_backfill.py`; credentialed copies under `/studio/media/credentialed/`). The worker now returns `render_params` (effective settings) with each output.

*Consent gate* on `POST /studio/jobs` (before the prompt policy): `kind:"voice"` → 422 `{error, category:"voice_off", consent?}` always (voice cloning is off). `params.likeness:true` or `params.consent_id` → needs an active consent whose scope covers the kind, else 422 `{category:"consent_required"}`. The receipt and credential then link the consent.

*Endpoints* (no token unless stated):
- `GET /provenance/status` → `{receipts, receipt_pubkey, c2pa, c2pa_library, c2pa_issuer:{subject, issuer, not_after, root_sha256, self_issued_root}, c2pa_trust, c2pa_note, watermark:{enabled, scheme, kinds}, voice_cloning:"off", counts, weights_digests}`.
- `GET /provenance/signing-key` → `{receipt:{scheme, pubkey, key_id, domains, same_key_as:"/attest/signing-key"}, c2pa:{issuer, trust, note, dev_ca_pem}}`.
- `POST /provenance/check?name=<file name>` with the raw file as the body (any content type; up to `DECOSA_PROVENANCE_MAX_UPLOAD_MB`, 50; 60 per hour per address, `DECOSA_PROVENANCE_CHECKS_PER_HOUR`). The file is checked in memory and not stored. → `{sha256, bytes, mime, checked_at, credential:{present, validation_state, codes:{success, informational, failure}, untrusted_issuer, content_intact, content_failures, chains_to_service_ca, signature, claim_generator, title, assertions, actions}, pixel_sha256, watermark:{detected, receipt_id, scheme, in_registry, note?}, matched_by:[exact_file|embedded_credential|raw_output|delivered_before_credential|same_pixels|watermark], receipt:{id, found_in_registry, signature_ok, key_pinned, signer, embedded_matches_registry, kind, model, license, weights_digest, seed, params_sha256, prompt_sha256, output_sha256, rendered_at, stamped_at, source, determinism, notice, receipt_url}|null, consent?, verdict, summary:[str]}`. `verdict` ∈ `credentialed` | `matched_without_credential` | `watermark_only` | `tampered` | `invalid_receipt` | `unknown`. `unknown` never means "not AI-generated".
- `GET /provenance/receipts/{rr_id}` → `{id, receipt, verification:{ok, pubkey_pinned, detail}, credential_url, gateway:null, note, consent?}`.
- `GET /provenance/samples` → `[{id, file, url, description}]` (files under `/studio/media/provenance-samples/`).
- `POST /provenance/consent-check` `{kind, consent_id?, likeness?}` → `{allowed:true, consent|null, note}` or `{allowed:false, category, reason, consent?}`. Dry run, nothing rendered.
- `POST /provenance/consents` (API key `dk_…` or the admin secret; demo tokens get 401) `{subject_label (pseudonym), clip_sha256, statement_sha256?, scope:{likeness:["voice"|"face"], kinds:["voice"|"music"|"image"|"video"], note?}, expires_in_days (1–3650, default 365)}` → 201 `{record, status}`. Record domain `decosa.consent-record.v1`.
- `GET /provenance/consents/{cr_id}` → `{id, state:"active"|"expired"|"revoked"|"invalid", signature_ok, key_pinned, scope, subject_label, clip_sha256, granted_at, expires_at, revoked_at, record_sha256}`.
- `POST /provenance/consents/{cr_id}/revoke` (API key or admin) `{reason?}` → the status. Revocation is signed (domain `decosa.consent-revocation.v1`).

Measured on our server 2026-09-23 (CPU): stamping an image ≈ 0.56 s (watermark, pixel hash, C2PA), a song or video ≈ 4 ms. Checks: credentialed image 0.9 s on the first call, stripped image 85 ms, song or video under 10 ms. Watermark (TrustMark Q) on the 6 gallery images: right id after JPEG q80 and q50, 50% resize and 80% centre crop (6/6 each), lost after a 50% crop or a 5° rotation (0/6); 1 of 6 unmarked originals decoded a random id, which the check reports as an unknown id, not a match. Demo run: `scripts/provenance_demo.py <base> out.json`; recording in `docs/provenance-demo-recording.json`.

**Endpoint auditor (added 2026-09-24).** New vertical id `auditor` (`POST /demo/session {"vertical":"auditor"}`; `dk_` keys work too). Code in `decosa_api/verticals/auditor/`.
- `GET /audit/signing-key` (no token) → `{alg: "ed25519", pubkey, key_id, domain, canonical}`. A dedicated auditor key, a key file in the API's config directory or `DECOSA_AUDITOR_KEY`; `DECOSA_AUDITOR_CREATE_KEY=1` creates it on first start. Without a key the audit routes return 503.
- `GET /audit/targets` (no token) → `[{id, label, claimed_model, kind: "gateway"|"openai", note, expect, available}]`: our own demo endpoints (`hosted`, `direct`, `swap`, `quant`), availability cached 10 s.
- `GET /audit/references`, `GET /audit/references/{id}` (no token) → the signed reference fixtures (`data/refs/<id>.json` + `.sig`, Ed25519 over `"decosa.audit.reference.v1\n" + sha256(file)`).
- `POST /audit/runs` (token) → SSE. Body `{target}` or `{base_url, model, claimed_model?, api_key?, context_tokens? (0–32000, default 12000), extra_body?}`.
  - Custom URLs: https only, and every resolved address must be globally routable (no loopback, private, link-local, CGNAT or reserved), the request is pinned to the checked address (Host header + SNI), no redirects, bodies capped at 4 MB. `DECOSA_AUDIT_ALLOW_PRIVATE=1` lifts this for self-host; `DECOSA_AUDIT_ALLOW_CUSTOM=0` turns custom URLs off.
  - `api_key` is held in memory for the run only: never logged, stored, or written into the report (`target.key_supplied` says whether one was given).
  - Errors: 400 bad body or refused URL, 402 under 3,000 tokens of budget left, 403 token for another vertical, 404 unknown target, 409 token already running an audit, 429 auditor busy (`DECOSA_AUDIT_MAX_CONCURRENT`, default 2), 503 demo target down or no key.
  - Events: `ready` `{target, claimed_model, reference, suite}`; `lane` for `identity`, `quality`, `performance` with `data.checks: [{id, title, status: pass|warn|fail|skip, value, band, detail}]` and `data.progress` (latest event per lane is its full state); `receipt` for every probe through the gateway; `lane` `verdict` (`data.progress: "recheck"` while re-checking, then `data.report`); `done` `{summary: {verdict, report, report_url}}`; a final `budget` event. Generated tokens count against the budget.
- Verdicts: `pass`, `drift`, `fail`, `inconclusive`. Drift and fail are signed only if a second pass of the golden prompts and canaries agrees; otherwise `inconclusive`. The policy thresholds are inside every report (`policy`).
- `GET /audit/reports/{id}` (no token) → the signed report. `GET /audit/reports` → the 20 newest reports for demo targets (custom-endpoint reports are never listed). `POST /audit/verify {report}` → `{valid_signature, signed_by_this_auditor, auditor_pubkey}`.
- Report signature: Ed25519 over `"decosa.audit.report.v1\n" + json.dumps(report minus "signature", sort_keys=True, separators=(",", ":"), ensure_ascii=False)` UTF-8; no floats (measurements are decimal strings), the same canonical JSON as `attest.py`. `scripts/verify_audit_report.py` checks one offline.
- On the gateway route, `usage` is the gateway's metering estimate, not the engine's token count, so the token-count fingerprint is skipped there.
