Checklist
waitingItems covered so far.
A finished inspection report with a checklist, issues graded by severity and measurements, each quoting what the inspector said and when.
Built on: Live speech to text
Start recording or run a sample.
Items covered so far.
Each with a severity.
Values and units heard.
Waiting for the first update.
Structured report when you stop.
Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.
The hosted demo stops at the live lanes. A self-hosted box then runs these steps on the full recording:
Recorded sessions from the live system, replayed event by event.
Loading recordings
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)# Decosa Field reports: use the hosted API
You are adding Decosa's Field reports to this project. Decosa runs open models (Voxtral Mini 4B Realtime for speech,
Qwen3.8-27B for text) through the Decosa API. Every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.
- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.
## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
`Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "field"}` returns
`{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
`GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
`GET /studio/jobs/{id}`, `GET /receipts/{id}`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.
## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=field&token=<token>`
Client to server:
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms each (1600 samples, 3200 bytes);
- a text frame `{"type":"stop"}` when the speaker is done. The server then writes final lanes, sends `done` and closes.
Server to client, JSON text frames:
- `{"type":"ready","vertical":"field","models":{"asr":"...","llm":"..."}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true}` (partials have `final:false` and are replaced by the next update)
- `{"type":"lane","lane":"<lane id>","title":"<Title>","body":"<markdown or text>","data":{...optional...},"latency_ms":850}`
Each lane event replaces the previous content of that lane.
- `{"type":"receipt","id":"<completion id>","model":"qwen3.8-27b","gateway_sig":"<hex>","provider":"<provider id>"}`
- `{"type":"budget","seconds_audio_left":240,"llm_tokens_left":15000}`
- `{"type":"error","message":"..."}`
- `{"type":"done","summary":{...final artifact...}}`
Lanes for `field`:
- `checklist`: inspection or site-walk checklist, ticked as items are covered
- `issues`: issues found, each with a severity
- `measurements`: measurements heard, with units
- `report`: on stop: a structured report (JSON in `data`) plus markdown in `body`
Final artifact in `done.summary`: the structured report JSON plus markdown.
Close codes the server uses:
- `4401` bad or expired token: get a new session, then reconnect.
- `4409` token already has a live socket: close the other one, or get a new session.
- `4429` busy (live-session cap): wait, then retry. Back off at least 15 s.
- `4400` bad parameters, `4402` budget exhausted, `4403` vertical or origin mismatch: do not retry; fix the cause.
On any other drop before `done`, reconnect with backoff (0.5 s, 1 s, 2 s, ... up to 8 s) using the same token while it
is valid. Stop sending audio when `seconds_audio_left` reaches 0.
## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `Authorization: Bearer <token>` and JSON
`{"vertical":"field","script_id":"<script id>"}` streams the same event types over SSE (`text/event-stream`, one JSON
event per `data:` line). It runs a canned audio script through the real pipeline, so the output is live.
Script ids for `field`: `field-roof-inspection`, `field-electrical-panel`. `GET https://api.decosa.ai/demo/scripts` lists all scripts (no token).
## Receipts (public, read-only)
`GET https://api.decosa.ai/receipts/{id}` returns
`{"id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id", "pubkey", "sig"}, "gateway": {"pubkey", "sig"}, "proof": {"format", "verified"}, "checks": [{"name", "ok", "detail"}]}`.
## Patient and client data
The hosted API is a public demo. It must not receive PHI or any real patient or client information. For real
data, self-host (see the "Run it yourself" prompt) so audio and text never leave the site.
## Audio format
Convert a recording to the wire format with ffmpeg:
`ffmpeg -i input.wav -ar 16000 -ac 1 -f s16le input.pcm`
## Example: TypeScript (Node 22+, global fetch and WebSocket)
```ts
import { readFileSync } from "node:fs";
const BASE = "https://api.decosa.ai";
// Prefer your API key (dk_…); fall back to a short demo session.
let token = process.env.DECOSA_API_KEY;
if (!token) {
const session = await fetch(`${BASE}/demo/session`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ vertical: "field" }),
});
if (session.status === 429) throw new Error(`busy, retry after ${session.headers.get("Retry-After")} s`);
({ token } = await session.json());
}
const ws = new WebSocket(
`${BASE.replace(/^http/, "ws")}/ws/live?vertical=field&token=${encodeURIComponent(token)}`,
);
ws.binaryType = "arraybuffer";
ws.onmessage = (m) => {
const ev = JSON.parse(String(m.data));
if (ev.type === "transcript" && ev.final) console.log("heard:", ev.text);
if (ev.type === "lane") console.log(`[${ev.lane}]`, ev.body);
if (ev.type === "receipt") console.log("receipt:", `${BASE}/receipts/${ev.id}`);
if (ev.type === "done") { console.log("final:", ev.summary); ws.close(); }
};
ws.onopen = async () => {
const pcm = readFileSync("input.pcm"); // 16 kHz mono PCM16 LE
for (let i = 0; i < pcm.length; i += 3200) {
ws.send(pcm.subarray(i, i + 3200));
await new Promise((r) => setTimeout(r, 100)); // real time
}
ws.send(JSON.stringify({ type: "stop" }));
};
```
## Example: Python (`pip install requests websockets`)
```python
import asyncio, json, os, requests, websockets
BASE = "https://api.decosa.ai"
token = os.environ.get("DECOSA_API_KEY") # dk_… API key preferred
if not token:
r = requests.post(f"{BASE}/demo/session", json={"vertical": "field"})
if r.status_code == 429:
raise SystemExit(f"busy, retry after {r.headers.get('Retry-After')} s")
token = r.json()["token"]
async def main():
url = BASE.replace("http", "ws", 1) + f"/ws/live?vertical=field&token={token}"
async with websockets.connect(url) as ws:
async def send_audio():
with open("input.pcm", "rb") as f: # 16 kHz mono PCM16 LE
while chunk := f.read(3200):
await ws.send(chunk)
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "stop"}))
sender = asyncio.create_task(send_audio())
async for msg in ws:
ev = json.loads(msg)
if ev["type"] == "lane":
print(f"[{ev['lane']}]", ev["body"])
elif ev["type"] == "receipt":
print("receipt:", f"{BASE}/receipts/{ev['id']}")
elif ev["type"] == "done":
print("final:", ev["summary"])
break
await sender
asyncio.run(main())
```
## What to build
1. A small client module for the calls above (session, socket, reconnect, 429 handling).
2. UI or CLI output that shows the transcript, each lane, and a link to every receipt.
3. A config value for the base URL, so it can point at a self-hosted box later with no code change.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Field reports: run it yourself (containers)
You are setting up Decosa Field reports to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
`curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
`"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"field"}'`
should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
## Sensitive data
Site walks can capture addresses, names and client details. Keep the service on a private network, with no public
port forwarding or tunnels.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/field-mac.md instead.
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Voxtral Mini 4B Realtime needs a GPU.
Needs about 36 GB of GPU memory at the smallest settings; 24 GB available.
Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.
The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.
Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available.
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
The standard tier fits (81.6 of 96 GB).
The standard tier fits (81.6 of 192 GB).
The standard tier fits (48 of 96 GB).
The standard tier fits (48 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa
curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yamlThe first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz
# {"ok": true, "asr": true, "llm": true, ...}
curl -fsS -X POST http://localhost:<PORT>/demo/session \
-H 'Content-Type: application/json' -d '{"vertical":"field"}'expected.json. Every check must print PASS.docker compose exec api python scripts/rehearse.py field
Download the mock-data bundle (665 KB, 9 checks)expected.json
The first 29 seconds of a synthetic roof inspection walk-through (12 Oak Street: granule loss and three missing shingles on the south slope, chimney step flashing lifted about two inches), streamed over the live WebSocket. The session must turn it into a structured inspection report with issues and measurements, each checked against what was said.
Licence: Synthetic: a script written for Decosa (no real people, patients or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-field-demo. Part of decosa-api, AGPL-3.0-or-later.
# Decosa Field reports: run it yourself (containers)
You are setting up Decosa Field reports to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
`curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
`"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"field"}'`
should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.
## Sensitive data
Site walks can capture addresses, names and client details. Keep the service on a private network, with no public
port forwarding or tunnels.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/field-mac.md instead.
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitField reports on GeForce RTX 5090
Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.
What this tool's stack says about this hardware:
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
The self-host prompt for Field reports, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Field reports on my hardware Fetch https://decosa.ai/prompts/field-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=field) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · 4-bit on a 32 GB card, plus a small card for speech (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Speech recognition: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB - Lanes and report: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB Warning: the fit check says this tier does not fit: Needs about 44 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/field-assemble.md
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live
# Decosa Field reports: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Field reports on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py field` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Speech recognition (streaming) | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured | | Lanes and report (language model) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py field`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 26 s · ~$0.026 per run · 28 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
Measured cost to run: about $0.026 per inspection (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
The inspector narrates the walk-through into a phone or laptop. Speech is transcribed live, and every ~10 seconds the language model updates a checklist for the inspection type, an issue list graded critical, major, minor or info, the measurements heard with any spec the inspector states, and the checks that passed. Each item carries the inspector's words, quoted verbatim from the transcript, with the time they were said. On stop it writes a structured report (JSON plus markdown). A safety sweep then checks the transcript against common safety items for the inspection type and adds any the inspector mentioned but the report left out. Finally, a claim check tests every report item against the transcript lines it cites: unsupported items are removed and listed for review, overstated ones are corrected. The report is ready to export as JSON or PDF. Self-hosted, the whole stack runs on one GPU box on-site and keeps working without an internet connection once the weights are downloaded.
The inspector's voice goes from a phone or browser mic as 16 kHz audio into Voxtral Mini 4B Realtime (Apache-2.0), which streams a transcript to decosa-api running the field pack. Every ~10 seconds decosa-api asks Qwen3.8-27B (Apache-2.0) to update four lanes: checklist, issues with severity, measurements with any stated spec, and passed checks, each tied to a quoted, timestamped transcript line. On stop it writes the report, runs a safety sweep for mentioned-but-missing safety items and a claim check of every item against its cited lines, then emits the report lane, which is exported as JSON, markdown or PDF. All of this sits inside a boundary marked as staying on-site when self-hosted. Below, the receipt path: on the gateway route each LLM call gets a signed receipt, the gateway countersigns it; anyone can check a receipt at GET /receipts/{id}. On the default self-host route receipts are signed by the box's own key (attested) and nothing leaves the box. An optional provider agent, off by default, can serve network jobs from the same LLM.
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
4-bit on a 32 GB card, plus a small card for speech
Same models and weights as Standard, so same output quality; you give up context length (16k) and concurrency (1–2 walk-throughs at a time).
one 96 GB card (the hosted demo)
Live Voxtral transcript plus Qwen3.8-27B NVFP4 with MTP; about 4 walk-throughs at once, 64k context.
two 96 GB cards
DeepSeek-V4-Flash writes the lanes and report; a final MOSS-Transcribe-Diarize pass cleans the transcript. Needs the whole two-card box and a community vLLM build.
the largest open flash models
GLM-5.3-Flash or DeepSeek-V4.1-Flash writes the report, served by network providers; speech recognition stays on your machine. Not served yet.
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Speech recognition (streaming)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab) 4.43B · 24 GBProof: partial | LiteStandardBestWanted | 4.43B · 24 GB | Proof: partial | |
| ||||
Lanes and report (language model)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57.6 GBProof: strong | LiteStandard | 27.8B · 57.6 GB | Proof: strong | |
| ||||
Final transcript pass (speaker-attributed)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.91BNo proof yet | Best | 0.91B | No proof yet | |
| ||||
Lanes and report (larger model)DeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yet | Best | 284B (13B active) · 192 GB | No proof yet | |
| ||||
Lanes and report (network-hosted flash model)GLM-5.3-Flash or DeepSeek-V4.1-Flashzai-org/GLM-5.3-Flash | deepseek-ai/DeepSeek-V4.1-Flash 321B (GLM-5.3-Flash) / 763B incl. Engram tables (V4.1-Flash) (18B (GLM-5.3-Flash) / null (V4.1-Flash) active) · about 170 GB (estimate)No proof yet | Wanted | 321B (GLM-5.3-Flash) / 763B incl. Engram tables (V4.1-Flash) (18B (GLM-5.3-Flash) / null (V4.1-Flash) active) · about 170 GB (estimate) | No proof yet | |
| ||||
${DECOSA_REGISTRY}/decosa-api:0.1.0Field pack lane engine: /ws/live?vertical=field, /demo/replay, /receipts/{id}, /healthz. No GPU. Publishing soon; builds from docker/api.
${DECOSA_REGISTRY}/decosa-llm:0.1.0Qwen3.8-27B on vLLM, OpenAI-compatible, served as qwen3.8-27b. Internal to the compose network.
${DECOSA_REGISTRY}/decosa-asr:0.1.0Voxtral realtime WebSocket, served as voxtral-realtime. Internal to the compose network.
Compose defaults (LLM 0.60, ASR 0.25 of the card), about 4 simultaneous walk-throughs. The latencies below were measured on our server with the two models on separate cards of this type.
No NVFP4 on Hopper: LLM_MODEL=Qwen/Qwen3.8-27B-FP8, LLM_GPU_UTIL=0.62. Starting point, not measured.
FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2. A walk-through uses well under 16k tokens. Not measured.
A 4-bit Qwen3.8-27B fits on its own (NVFP4 weights 19.9 GiB; Q4 GGUF 17.5–21 GB), but not together with the ASR model and a useful KV cache on the same card.
Smaller-box option: LLM alone on the 32 GB card with a short context (16k) and 1–2 live sessions, Voxtral (8.3 GB BF16 weights) on the second card. Estimate from the lineup page; not measured.
Measuredmeasured on our server 2026-09-23 (decosa-api ops/record-all.json, field scripts: 800 and 2000 ms)
Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 3251 and 4448 ms (max 5887)
Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 2518 and 4448 ms (max 5887)
Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 3985 and 4448 ms (max 5887)
Measuredmeasured on our server 2026-09-23, direct route, 8 field walk-throughs 2 at a time: median 8.5 s (7.4–11.5); gateway route, one roof replay: 9.9 s
Measuredmeasured on our server 2026-09-23, direct route, n = 8: median 0.21 s (0.16–0.24)
Measuredmeasured on our server 2026-09-23, direct route, n = 8: median 2.3 s (2.0–4.6), one judge call per report item, 10 in parallel; gateway route, one roof replay: 4.2 s
Measuredmeasured on our server 2026-09-23, direct route, n = 8: done event a median 16.1 s (12.8–21.1) after the last transcript line; before the sweep and claim check it was 8.1 and 11.6 s on the gateway route
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa field-reports stack on this machine
You are setting up a self-hosted voice-to-inspection-report stack on this Linux machine. Work step by step, show me each command before you run anything that installs software, and stop to ask if a check fails. Everything runs locally: audio, transcripts and reports stay on this box, and it works offline once the weights are downloaded.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | model / role | licence | engine (pinned) | port |
|---|---|---|---|---|
| `llm` | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80` (27.8B dense) | Apache-2.0 | vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`, MTP 3 draft tokens, FP8 KV cache | internal 8000 |
| `asr` | `mistralai/Voxtral-Mini-4B-Realtime-2602` (4.43B, BF16) | Apache-2.0 | vLLM 0.27.1 (`vllm/vllm-openai:v0.27.1` + `mistral-common[audio]`), realtime WebSocket | internal 8000 |
| `api` | `decosa-api`: field pack (lanes `checklist`, `issues`, `measurements`, `passed_checks`; `safety_sweep`, `report_check` and `report` on stop) | – | Python 3.12, no GPU | `127.0.0.1:8445` |
Images: `${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0` (publishing soon). Until they are published, build them from the `decosa-api` source repo (`docker compose build`, Dockerfiles under `docker/`).
## 1. Check the GPU, driver and Docker
1. `nvidia-smi`: note the GPU model, memory and driver. The images were built for driver 580 or newer.
2. Pick the settings for this GPU:
- **Blackwell, 96 GB (RTX PRO 6000, B200):** the defaults below.
- **Hopper 80–141 GB (H100/H200):** no NVFP4. Set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
- **48 GB (L40S, RTX 6000 Ada):** FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`. A walk-through uses well under 16k tokens.
- **Smaller box, two cards:** one 32 GB Blackwell card for the NVFP4 LLM alone (`LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.90`, `DECOSA_LIVE_CAP=1`) and a second card with 16 GB or more for the ASR (`ASR_GPU_UTIL=0.80`). Pin each service to its own card (step 3). The 4-bit LLM needs about 20 GiB for weights, so a single 24–32 GB card cannot also hold the ASR. This layout has not been measured; treat it as a starting point.
- **Anything smaller:** stop and tell me; it will not fit both models.
3. `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt repository for this distro.
4. `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. If that fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit` from NVIDIA's apt repo), run `sudo nvidia-ctk runtime configure --runtime=docker`, restart Docker, and retry.
5. Check free disk: about 80 GB for images and weights (weights alone: about 21 GB for the LLM, 8.3 GB for the ASR).
## 2. Get the source
```bash
git clone <decosa-api repo URL> decosa && cd decosa
cp .env.example .env
```
Edit `.env` for the GPU (step 1). Keep `DECOSA_BIND=127.0.0.1` and `DECOSA_LLM_ROUTE=direct` so nothing leaves the box.
## 3. docker-compose.yml
Use the repo's `docker-compose.yml`. Check it against this and fix any drift:
```yaml
name: decosa
x-gpu: &gpu
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG:-0.1.0}
build: docker/llm
<<: *gpu
ipc: host
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL:-nvidia/Qwen3.8-27B-NVFP4}", "--revision", "${LLM_REVISION:-482ca0f3832238542f8f5295dde86b5f22711d80}",
"--served-model-name", "qwen3.8-27b", "--language-model-only",
"--max-model-len", "${LLM_MAX_LEN:-65536}", "--gpu-memory-utilization", "${LLM_GPU_UTIL:-0.60}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3",
"--speculative-config", '{"method":"mtp","num_speculative_tokens":3}', "--seed", "0",
"--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 5, start_period: 900s }
asr:
image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG:-0.1.0}
build: docker/asr
<<: *gpu
ipc: host
depends_on: { llm: { condition: service_healthy } }
volumes: [hf-cache:/root/.cache/huggingface]
command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
"--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
"--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL:-0.25}",
"--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 5, start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG:-0.1.0}
build: { context: ., dockerfile: docker/api/Dockerfile }
depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
environment:
DECOSA_ASR_WS: ws://asr:8000/v1/realtime
DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE:-direct}
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP:-4}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_AUDIO_S: "3600"
DECOSA_BUDGET_LLM_TOKENS: "200000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_CORS_ORIGINS: ${DECOSA_CORS_ORIGINS:-}
ports: ["${DECOSA_BIND:-127.0.0.1}:${DECOSA_PORT:-8445}:8445"]
volumes: [decosa-data:/data]
volumes: { hf-cache: {}, decosa-data: {} }
```
For the two-card layout, give `asr` its own device: replace `<<: *gpu` in `asr` with the same block using `device_ids: ["1"]`. The shares must add up to under about 0.9 when both engines share one card.
## 4. Start and wait
```bash
docker compose pull || docker compose build # images are "publishing soon"; build falls back to source
docker compose up -d
docker compose ps # llm healthy first (5-10 min on first start), then asr (1-3 min)
curl -s localhost:8445/healthz # expect "ok": true, "asr": true, "llm": true, "llm_route": "direct"
```
If `llm` stays unhealthy, read `docker compose logs llm`: usually out of memory (lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`) or NVFP4 on a pre-Blackwell GPU (switch to the FP8 checkpoint).
## 5. Smoke test: replay a sample walk-through
The replay route runs a canned 16 kHz recording through the real ASR and LLM pipeline and streams the same events as the live WebSocket, over SSE.
```bash
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"field"}' | jq -r .token)
curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"vertical":"field","script_id":"field-roof-inspection"}' | tee roof.sse | grep -o '"type": *"[a-z]*"' | sort | uniq -c
```
Pass when the stream has a `ready` event, `transcript` events, `lane` events for `checklist`, `issues`, `measurements` and `passed_checks` during the run, then one `lane` event each for `safety_sweep`, `report_check` and `report` after the audio ends, and a final `done`. Expect `receipt` events with `"status": "attested"` on the direct route; signed by the box's own key rather than a gateway. The other sample is `field-electrical-panel`. Lane events replace the lane's previous state; they do not append.
## 6. Export the report (JSON, markdown, PDF)
The `done` event carries `summary.report` (structured JSON: title, summary, overall_condition, issues with severity and priority, measurements with the stated spec and within_spec, passed_checks, code_references, site_details, checklist, recommendations, follow_up; every issue, measurement, passed check, code reference and site detail carries `quote` (the inspector's words, verbatim), `lines`, `t` (seconds into the walk-through) and `verified`) and `summary.report_markdown`.
```bash
grep '^data: ' roof.sse | sed 's/^data: //' | jq -c 'select(.type=="done") | .summary' > done.json
jq '.report' done.json > roof-report.json
jq -r '.report_markdown' done.json > roof-report.md
pandoc roof-report.md -o roof-report.pdf # optional; needs pandoc and a PDF engine
jq -r '.issues[] | "\(.severity)\t\(.issue)\t\(.location)"' roof-report.json
```
## 7. Point the app at the local API
- Set the web app's API base to `http://localhost:8445` (for the Decosa web app, `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`).
- Live mic: open `ws://localhost:8445/ws/live?vertical=field&token=$TOKEN`, send 16 kHz mono PCM16 little-endian frames of about 100 ms, then the text frame `{"type":"stop"}`, and read the JSON events.
- If the page is served from another origin, add it to `DECOSA_CORS_ORIGINS` (a WebSocket close with code `4403` means the origin was refused).
- To reach it from phones on site, put a TLS reverse proxy on the LAN in front of port 8445 and set `DECOSA_TRUSTED_PROXIES`. Do not expose it to the internet.
When done, report the GPU tier you chose, the `/healthz` output, the lane ids seen in the smoke test, and the path of the exported report files.Last reviewed
With Decosa field reports, the inspector narrates the walk-through into a phone or laptop. Speech is transcribed live, and about every 10 seconds a language model updates a checklist, an issue list graded critical, major, minor or info, measurements heard, and passed checks, each quoted from the transcript with its time. On stop it writes a JSON and Markdown report, runs a safety sweep and a claim check. The report is a draft for the inspector to review and sign.
On Decosa's internal eval, field reports included 40/41 (98%) expected issues, rated 40/40 found issues within the accepted severity range with all 6 critical-only hazards rated critical, and kept 51/51 stated facts; the judge flagged 5/125 claims. The caveats matter: 8 synthetic TTS walk-throughs, no real site recordings, prompts tuned on these same scripts so there is no held-out set, and a Gemma 4 judge checked by hand only in part.
Decosa field reports can be self-hosted: the whole stack runs on one GPU box on-site and keeps working without an internet connection once the model weights are downloaded. With DECOSA_LLM_ROUTE=direct, nothing leaves the box. A lite tier runs on a 32 GB Blackwell card plus a card of 16 GB or more for speech, with 16k context and 1 to 2 walk-throughs at a time.
No. Decosa field reports check each measurement against the spec the inspector states aloud; there is no built-in code database. The report is a drafting aid for the inspector to review and sign, and it does not replace a licensed inspection or a code-compliance determination. Self-hosting keeps site audio, transcripts and reports on your own hardware.
Decosa field reports write the final report after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load (hosted p50 of 26 s over 3 replays). That 75 s roof sample cost about $0.026 at the gateway list price, about 41k tokens and 28 model calls. Output is structured JSON and Markdown, with PDF via pandoc.
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…