Skip to content
decosa
LabsHostedSelf-hostMac

Write the inspection report

A finished inspection report with a checklist, issues graded by severity and measurements, each quoting what the inspector said and when.

Measured40/41 (98%); before 38/41 (93%)Expected issues in the final report, now vs before (synthetic set)
On production26 smedian on production (2026-09-25); slower when the service is busy
List price~$0.026 per inspectionmeasured, at list price

Built on: Live speech to text

Try it live

Live
Tick the consent box to record.

Demo audio left05:00
Demo usage left100%

Live transcript · Voxtral Mini 4B Realtime

Start recording or run a sample.

Live lanes

Checklist

waiting

Items covered so far.

Issues

waiting

Each with a severity.

Measurements

waiting

Values and units heard.

Passed checks

waiting

Waiting for the first update.

Inspection report

waiting

Structured report when you stop.

Receipts

Proof · signed records appear as each step finishes

Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.

After the session, on your own GPU

Self-host only

The hosted demo stops at the live lanes. A self-hosted box then runs these steps on the full recording:

  1. 01Safety sweep
  2. 02Claim check
Watch a recorded run first

Watch a recorded session

Live

Recorded sessions from the live system, replayed event by event.

Loading recordings

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the field reports API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
field

Use the hosted API

# Decosa Field reports: use the hosted API

You are adding Decosa's Field reports to this project. Decosa runs open models (Voxtral Mini 4B Realtime for speech,
Qwen3.8-27B for text) through the Decosa API. Every model output comes with a signed receipt.
Use only the endpoints below. If you need something that is not listed, stop and ask me; do not guess endpoints.

- Base URL: `https://api.decosa.ai`
- WebSocket base: `wss://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.

## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
   only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
   `Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "field"}` returns
   `{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
   render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
   `GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
   `GET /studio/jobs/{id}`, `GET /receipts/{id}`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
   HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.

## Live audio
`WS wss://api.decosa.ai/ws/live?vertical=field&token=<token>`

Client to server:
- binary frames: 16 kHz mono PCM16 little-endian, about 100 ms each (1600 samples, 3200 bytes);
- a text frame `{"type":"stop"}` when the speaker is done. The server then writes final lanes, sends `done` and closes.

Server to client, JSON text frames:
- `{"type":"ready","vertical":"field","models":{"asr":"...","llm":"..."}}`
- `{"type":"transcript","t":12.4,"text":"...","final":true}` (partials have `final:false` and are replaced by the next update)
- `{"type":"lane","lane":"<lane id>","title":"<Title>","body":"<markdown or text>","data":{...optional...},"latency_ms":850}`
  Each lane event replaces the previous content of that lane.
- `{"type":"receipt","id":"<completion id>","model":"qwen3.8-27b","gateway_sig":"<hex>","provider":"<provider id>"}`
  
- `{"type":"budget","seconds_audio_left":240,"llm_tokens_left":15000}`
- `{"type":"error","message":"..."}`
- `{"type":"done","summary":{...final artifact...}}`

Lanes for `field`:
- `checklist`: inspection or site-walk checklist, ticked as items are covered
- `issues`: issues found, each with a severity
- `measurements`: measurements heard, with units
- `report`: on stop: a structured report (JSON in `data`) plus markdown in `body`

Final artifact in `done.summary`: the structured report JSON plus markdown.

Close codes the server uses:
- `4401` bad or expired token: get a new session, then reconnect.
- `4409` token already has a live socket: close the other one, or get a new session.
- `4429` busy (live-session cap): wait, then retry. Back off at least 15 s.
- `4400` bad parameters, `4402` budget exhausted, `4403` vertical or origin mismatch: do not retry; fix the cause.

On any other drop before `done`, reconnect with backoff (0.5 s, 1 s, 2 s, ... up to 8 s) using the same token while it
is valid. Stop sending audio when `seconds_audio_left` reaches 0.

## Testing without a microphone
`POST https://api.decosa.ai/demo/replay` with `Authorization: Bearer <token>` and JSON
`{"vertical":"field","script_id":"<script id>"}` streams the same event types over SSE (`text/event-stream`, one JSON
event per `data:` line). It runs a canned audio script through the real pipeline, so the output is live.
Script ids for `field`: `field-roof-inspection`, `field-electrical-panel`. `GET https://api.decosa.ai/demo/scripts` lists all scripts (no token).

## Receipts (public, read-only)
`GET https://api.decosa.ai/receipts/{id}` returns
`{"id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id", "pubkey", "sig"}, "gateway": {"pubkey", "sig"}, "proof": {"format", "verified"}, "checks": [{"name", "ok", "detail"}]}`.

## Patient and client data
The hosted API is a public demo. It must not receive PHI or any real patient or client information. For real
data, self-host (see the "Run it yourself" prompt) so audio and text never leave the site.

## Audio format
Convert a recording to the wire format with ffmpeg:
`ffmpeg -i input.wav -ar 16000 -ac 1 -f s16le input.pcm`

## Example: TypeScript (Node 22+, global fetch and WebSocket)
```ts
import { readFileSync } from "node:fs";

const BASE = "https://api.decosa.ai";
// Prefer your API key (dk_…); fall back to a short demo session.
let token = process.env.DECOSA_API_KEY;
if (!token) {
  const session = await fetch(`${BASE}/demo/session`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ vertical: "field" }),
  });
  if (session.status === 429) throw new Error(`busy, retry after ${session.headers.get("Retry-After")} s`);
  ({ token } = await session.json());
}

const ws = new WebSocket(
  `${BASE.replace(/^http/, "ws")}/ws/live?vertical=field&token=${encodeURIComponent(token)}`,
);
ws.binaryType = "arraybuffer";
ws.onmessage = (m) => {
  const ev = JSON.parse(String(m.data));
  if (ev.type === "transcript" && ev.final) console.log("heard:", ev.text);
  if (ev.type === "lane") console.log(`[${ev.lane}]`, ev.body);
  if (ev.type === "receipt") console.log("receipt:", `${BASE}/receipts/${ev.id}`);
  if (ev.type === "done") { console.log("final:", ev.summary); ws.close(); }
};
ws.onopen = async () => {
  const pcm = readFileSync("input.pcm"); // 16 kHz mono PCM16 LE
  for (let i = 0; i < pcm.length; i += 3200) {
    ws.send(pcm.subarray(i, i + 3200));
    await new Promise((r) => setTimeout(r, 100)); // real time
  }
  ws.send(JSON.stringify({ type: "stop" }));
};
```

## Example: Python (`pip install requests websockets`)
```python
import asyncio, json, os, requests, websockets

BASE = "https://api.decosa.ai"
token = os.environ.get("DECOSA_API_KEY")  # dk_… API key preferred
if not token:
    r = requests.post(f"{BASE}/demo/session", json={"vertical": "field"})
    if r.status_code == 429:
        raise SystemExit(f"busy, retry after {r.headers.get('Retry-After')} s")
    token = r.json()["token"]

async def main():
    url = BASE.replace("http", "ws", 1) + f"/ws/live?vertical=field&token={token}"
    async with websockets.connect(url) as ws:
        async def send_audio():
            with open("input.pcm", "rb") as f:  # 16 kHz mono PCM16 LE
                while chunk := f.read(3200):
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)
            await ws.send(json.dumps({"type": "stop"}))
        sender = asyncio.create_task(send_audio())
        async for msg in ws:
            ev = json.loads(msg)
            if ev["type"] == "lane":
                print(f"[{ev['lane']}]", ev["body"])
            elif ev["type"] == "receipt":
                print("receipt:", f"{BASE}/receipts/{ev['id']}")
            elif ev["type"] == "done":
                print("final:", ev["summary"])
                break
        await sender

asyncio.run(main())
```

## What to build
1. A small client module for the calls above (session, socket, reconnect, 429 handling).
2. UI or CLI output that shows the transcript, each lane, and a link to every receipt.
3. A config value for the base URL, so it can point at a self-hosted box later with no code change.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Field reports: run it yourself (containers)

You are setting up Decosa Field reports to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"field"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Sensitive data
Site walks can capture addresses, names and client details. Keep the service on a private network, with no public
port forwarding or tunnels.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/field-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Voxtral Mini 4B Realtime needs a GPU.

  • GeForce RTX 4090Doesn't fit

    Needs about 36 GB of GPU memory at the smallest settings; 24 GB available.

  • GeForce RTX 5090Doesn't fit

    Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.

  • L40SDoesn't fit

    Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (81.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (81.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (48 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (48 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"field"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py field

Download the mock-data bundle (665 KB, 9 checks)expected.json

The first 29 seconds of a synthetic roof inspection walk-through (12 Oak Street: granule loss and three missing shingles on the south slope, chimney step flashing lifted about two inches), streamed over the live WebSocket. The session must turn it into a structured inspection report with issues and measurements, each checked against what was said.

What the rehearsal checks
  • the session ends with a done event
  • the session reports no errors
  • the report is titled for the site (Oak Street)
  • at least 2 issues are listed
  • the missing shingles and the lifted chimney flashing are among the issues
  • the roof pitch is recorded as a measurement
  • the overall condition is fair, poor or unsafe (not good)
  • at least 3 report claims are checked and supported by the transcript
  • every speech and model call has a signed receipt

Licence: Synthetic: a script written for Decosa (no real people, patients or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-field-demo. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Field reports: run it yourself (containers)

You are setting up Decosa Field reports to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"field"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

## Sensitive data
Site walks can capture addresses, names and client details. Keep the service on a private network, with no public
port forwarding or tunnels.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/field-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitField reports on GeForce RTX 5090

Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.

What this tool's stack says about this hardware:

  • 2 cards: 32 GB Blackwell (NVFP4 LLM) + ≥16 GB (ASR) (fits): Smaller-box option: LLM alone on the 32 GB card with a short context (16k) and 1–2 live sessions, Voxtral (8.3 GB BF16 weights) on the second card. Estimate from the lineup page; not measured.

Lite · 4-bit on a 32 GB card, plus a small card for speech: what changesuses estimates

  • Needs about 44 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • Speech recognition: Voxtral Mini 4B Realtime. ~24 GB (at least ~16 GB), weights 8.3 GB (from stack.json). Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more.
  • Lanes and report: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Field reports, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Field reports on my hardware

Fetch https://decosa.ai/prompts/field-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=field)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · 4-bit on a 32 GB card, plus a small card for speech (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Speech recognition: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB
- Lanes and report: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB

Warning: the fit check says this tier does not fit: Needs about 44 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/field-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live

Mac prompt for your coding agent

# Decosa Field reports: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Field reports on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py field` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Speech recognition (streaming) | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured |
| Lanes and report (language model) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py field`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 26 s · ~$0.026 per run · 28 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

Measured cost to run: about $0.026 per inspection (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Known limits (2)
  • The report is written after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load.
  • Measurements are checked against the spec the inspector states; there is no built-in code database.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the field reports API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Talk through an inspection, site visit or job; get the checklist, issues by severity, measurements and a finished report.

The inspector narrates the walk-through into a phone or laptop. Speech is transcribed live, and every ~10 seconds the language model updates a checklist for the inspection type, an issue list graded critical, major, minor or info, the measurements heard with any spec the inspector states, and the checks that passed. Each item carries the inspector's words, quoted verbatim from the transcript, with the time they were said. On stop it writes a structured report (JSON plus markdown). A safety sweep then checks the transcript against common safety items for the inspection type and adds any the inspector mentioned but the report left out. Finally, a claim check tests every report item against the transcript lines it cites: unsupported items are removed and listed for review, overstated ones are corrected. The report is ready to export as JSON or PDF. Self-hosted, the whole stack runs on one GPU box on-site and keeps working without an internet connection once the weights are downloaded.

Deployment
Hosted or self-host
Regulatory
Drafting aid: the report is a draft for the inspector to review and sign; it does not replace a licensed inspection or code-compliance determination. Self-host keeps site audio, transcripts and reports on your hardware.
Architecture
Text description

The inspector's voice goes from a phone or browser mic as 16 kHz audio into Voxtral Mini 4B Realtime (Apache-2.0), which streams a transcript to decosa-api running the field pack. Every ~10 seconds decosa-api asks Qwen3.8-27B (Apache-2.0) to update four lanes: checklist, issues with severity, measurements with any stated spec, and passed checks, each tied to a quoted, timestamped transcript line. On stop it writes the report, runs a safety sweep for mentioned-but-missing safety items and a claim check of every item against its cited lines, then emits the report lane, which is exported as JSON, markdown or PDF. All of this sits inside a boundary marked as staying on-site when self-hosted. Below, the receipt path: on the gateway route each LLM call gets a signed receipt, the gateway countersigns it; anyone can check a receipt at GET /receipts/{id}. On the default self-host route receipts are signed by the box's own key (attested) and nothing leaves the box. An optional provider agent, off by default, can serve network jobs from the same LLM.

Architecture

At a glance

Measured latency
Hosted p50 below is the time from the end of the audio to the final `done`, over 3 replays of the sample (2026-09-25, gateway route, one session at a time).
Typical run cost
A few cents for the roof sample (a few dozen model calls) at the gateway list price. Each run shows its own measured cost.
Data retention
Audio and transcript in memory for the session only; the report comes back in done and is not stored.
What leaves the box (self-host)
Nothing, with DECOSA_LLM_ROUTE=direct.
Outputs
Structured JSON (issues with severity, quote, line and time; measurements; passed checks) and Markdown; PDF with pandoc.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    4-bit on a 32 GB card, plus a small card for speech

    Same models and weights as Standard, so same output quality; you give up context length (16k) and concurrency (1–2 walk-throughs at a time).

    Models
    • Voxtral Mini 4B Realtime
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1× 32 GB Blackwell card (NVFP4 LLM) + 1× ≥16 GB card (Voxtral)
    Quality evidence
    • Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM478coding-agent-bench README (our server)
    • Same weights served by llama.cpp (Q4 GGUF path): avoid for this tier62–63 / 500coding-agent-bench README (our server)
    • Field lane qualitynot measured yet
    Latency
    estimate: similar per-call latency to Standard at 1 session; not measured on a 32 GB card
    Verification
    Proof: strongSelf-host onlyQwen3.8-27B is the hosted model qwen3.8-27b: gateway-signed receipts on the gateway route; the default self-host route (direct) signs receipts with the box's own key (attested).
  • In the hosted demo

    Standard

    one 96 GB card (the hosted demo)

    Live Voxtral transcript plus Qwen3.8-27B NVFP4 with MTP; about 4 walk-throughs at once, 64k context.

    Models
    • Voxtral Mini 4B Realtime
    • Qwen3.8-27B (NVIDIA NVFP4)
    Hardware
    1× RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM478coding-agent-bench README (our server)
    • ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; FP8 on vLLM)34.2scribe-bench wiki/models.md
    • Quality smoke, math / code38/40, 12/12Decosa model benchmarks (Sep 2026)
    • Field report, internal synthetic eval (n = 8 walk-throughs: 2 recorded demo + 6 new), with the grounded report, safety sweep and claim check: expected issues in the report40/41 (98%); before 38/41 (93%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
    • Same eval: severity within the accepted range, of issues found40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
    • Same eval: stated facts kept (incl. passed checks and spec limits, now fields of their own)51/51; before 32/51 (63%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
    • Same eval: missed safety item (bathroom receptacles without GFCI)in the report in 3 of 3 re-runs. Run on the original report, the safety sweep adds it as critical, quoted at 1:03eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
    • Same eval: claims the judge flagged as contradicted or unsupported5/125 (before 0/100, plus 1 found by hand: 'cracked chimney' where the inspector said sealant). By hand: no invented findings. 2 of the 5 are the report's overall-condition rating; 3 are speech-recognition spellings now quoted verbatim ('ceiling' for sealant, 'pole' for pull stations, 'stab block' for Stab-Lok), which need an ASR glossary, not a model changeeval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
    Latency
    measured on our server: lanes in seconds; on stop, the report, safety sweep and claim check (direct route) bring the finished report seconds after the audio ends
    Verification
    Proof: strongQwen3.8-27B is the hosted model qwen3.8-27b: gateway-signed receipts on the gateway route; the default self-host route (direct) signs receipts with the box's own key (attested).
  • Best

    two 96 GB cards

    DeepSeek-V4-Flash writes the lanes and report; a final MOSS-Transcribe-Diarize pass cleans the transcript. Needs the whole two-card box and a community vLLM build.

    Models
    • Voxtral Mini 4B Realtime
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2× RTX PRO 6000 Blackwell 96 GB (TP2)
    Quality evidence
    • ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; run on a Mac via MLX, not the NVFP4 kit)35.8scribe-bench wiki/models.md
    • Medical-term miss rate on PriMock57, MOSS-Transcribe-Diarize vs Nemotron-3.5 streaming (%)8.4 vs 12.7scribe-bench RESULTS.md / wiki/decoder-finding.md
    • Owner's coding benchmark, core tasks (/500), DeepSeek V4-Flash453–459coding-agent-bench README (our server)
    • Field lane qualitynot measured yet
    Latency
    measured single-stream decode 108.8 → 150.6 tok/s with MTP (dsv4-flash-nvfp4-sm120 README); field lane latency not measured
    Verification
    No proof yetSelf-host onlyNo receipts today: neither DeepSeek-V4-Flash nor the MOSS pass is a hosted model. Only the live Voxtral step is shared with Standard.
  • Needs more compute

    Wanted: the best setup

    the largest open flash models

    GLM-5.3-Flash or DeepSeek-V4.1-Flash writes the report, served by network providers; speech recognition stays on your machine. Not served yet.

    Models
    • Voxtral Mini 4B Realtime
    • GLM-5.3-Flash or DeepSeek-V4.1-Flash
    Hardware
    Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). Speech recognition stays local. Estimate.
    Quality evidence
    • Field lane qualitynot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Speech recognition (streaming)Voxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab)
4.43B · 24 GBProof: partial
Lanes and report (language model)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57.6 GBProof: strong
Final transcript pass (speaker-attributed)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.91BNo proof yet
Lanes and report (larger model)DeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yet
Lanes and report (network-hosted flash model)GLM-5.3-Flash or DeepSeek-V4.1-Flashzai-org/GLM-5.3-Flash | deepseek-ai/DeepSeek-V4.1-Flash
321B (GLM-5.3-Flash) / 763B incl. Engram tables (V4.1-Flash) (18B (GLM-5.3-Flash) / null (V4.1-Flash) active) · about 170 GB (estimate)No proof yet

Around the models

Tools, services and hardware

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Field pack lane engine: /ws/live?vertical=field, /demo/replay, /receipts/{id}, /healthz. No GPU. Publishing soon; builds from docker/api.

  • llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    Qwen3.8-27B on vLLM, OpenAI-compatible, served as qwen3.8-27b. Internal to the compose network.

  • asr:8000
    ${DECOSA_REGISTRY}/decosa-asr:0.1.0

    Voxtral realtime WebSocket, served as voxtral-realtime. Internal to the compose network.

Hardware

  • 1× RTX PRO 6000 Blackwell 96 GB Fits

    Compose defaults (LLM 0.60, ASR 0.25 of the card), about 4 simultaneous walk-throughs. The latencies below were measured on our server with the two models on separate cards of this type.

  • 1× H100 / H200 (80–141 GB) Fits

    No NVFP4 on Hopper: LLM_MODEL=Qwen/Qwen3.8-27B-FP8, LLM_GPU_UTIL=0.62. Starting point, not measured.

  • 1× L40S / RTX 6000 Ada 48 GB Fits

    FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2. A walk-through uses well under 16k tokens. Not measured.

  • 1× 24–32 GB card Does not fit

    A 4-bit Qwen3.8-27B fits on its own (NVFP4 weights 19.9 GiB; Q4 GGUF 17.5–21 GB), but not together with the ASR model and a useful KV cache on the same card.

  • 2 cards: 32 GB Blackwell (NVFP4 LLM) + ≥16 GB (ASR) Fits

    Smaller-box option: LLM alone on the 32 GB card with a short context (16k) and 1–2 live sessions, Voxtral (8.3 GB BF16 weights) on the second card. Estimate from the lineup page; not measured.

Latency per lane

  • transcript (first text)1.4 s

    Measuredmeasured on our server 2026-09-23 (decosa-api ops/record-all.json, field scripts: 800 and 2000 ms)

  • checklist3.9 s

    Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 3251 and 4448 ms (max 5887)

  • issues3.5 s

    Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 2518 and 4448 ms (max 5887)

  • measurements4.2 s

    Measuredmeasured on our server 2026-09-23, gateway route; mean of per-script medians 3985 and 4448 ms (max 5887)

  • report (writer)8.5 s

    Measuredmeasured on our server 2026-09-23, direct route, 8 field walk-throughs 2 at a time: median 8.5 s (7.4–11.5); gateway route, one roof replay: 9.9 s

  • safety_sweep210 ms

    Measuredmeasured on our server 2026-09-23, direct route, n = 8: median 0.21 s (0.16–0.24)

  • report_check2.3 s

    Measuredmeasured on our server 2026-09-23, direct route, n = 8: median 2.3 s (2.0–4.6), one judge call per report item, 10 in parallel; gateway route, one roof replay: 4.2 s

  • report ready after audio ends16.1 s

    Measuredmeasured on our server 2026-09-23, direct route, n = 8: done event a median 16.1 s (12.8–21.1) after the last transcript line; before the sweep and claim check it was 8.1 and 11.6 s on the gateway route

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

field/assemble-prompt.md159 lines
# Assemble the Decosa field-reports stack on this machine

You are setting up a self-hosted voice-to-inspection-report stack on this Linux machine. Work step by step, show me each command before you run anything that installs software, and stop to ask if a check fails. Everything runs locally: audio, transcripts and reports stay on this box, and it works offline once the weights are downloaded.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/field.zip (665 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py field` (the api image carries the same bundle under /app/rehearsal/field/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the session ends with a done event", "the session reports no errors", "the report is titled for the site (Oak Street)"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | model / role | licence | engine (pinned) | port |
|---|---|---|---|---|
| `llm` | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80` (27.8B dense) | Apache-2.0 | vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`, MTP 3 draft tokens, FP8 KV cache | internal 8000 |
| `asr` | `mistralai/Voxtral-Mini-4B-Realtime-2602` (4.43B, BF16) | Apache-2.0 | vLLM 0.27.1 (`vllm/vllm-openai:v0.27.1` + `mistral-common[audio]`), realtime WebSocket | internal 8000 |
| `api` | `decosa-api`: field pack (lanes `checklist`, `issues`, `measurements`, `passed_checks`; `safety_sweep`, `report_check` and `report` on stop) | – | Python 3.12, no GPU | `127.0.0.1:8445` |

Images: `${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0` (publishing soon). Until they are published, build them from the `decosa-api` source repo (`docker compose build`, Dockerfiles under `docker/`).

## 1. Check the GPU, driver and Docker

1. `nvidia-smi`: note the GPU model, memory and driver. The images were built for driver 580 or newer.
2. Pick the settings for this GPU:
   - **Blackwell, 96 GB (RTX PRO 6000, B200):** the defaults below.
   - **Hopper 80–141 GB (H100/H200):** no NVFP4. Set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
   - **48 GB (L40S, RTX 6000 Ada):** FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`. A walk-through uses well under 16k tokens.
   - **Smaller box, two cards:** one 32 GB Blackwell card for the NVFP4 LLM alone (`LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.90`, `DECOSA_LIVE_CAP=1`) and a second card with 16 GB or more for the ASR (`ASR_GPU_UTIL=0.80`). Pin each service to its own card (step 3). The 4-bit LLM needs about 20 GiB for weights, so a single 24–32 GB card cannot also hold the ASR. This layout has not been measured; treat it as a starting point.
   - **Anything smaller:** stop and tell me; it will not fit both models.
3. `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt repository for this distro.
4. `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. If that fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit` from NVIDIA's apt repo), run `sudo nvidia-ctk runtime configure --runtime=docker`, restart Docker, and retry.
5. Check free disk: about 80 GB for images and weights (weights alone: about 21 GB for the LLM, 8.3 GB for the ASR).

## 2. Get the source

```bash
git clone <decosa-api repo URL> decosa && cd decosa
cp .env.example .env
```

Edit `.env` for the GPU (step 1). Keep `DECOSA_BIND=127.0.0.1` and `DECOSA_LLM_ROUTE=direct` so nothing leaves the box.

## 3. docker-compose.yml

Use the repo's `docker-compose.yml`. Check it against this and fix any drift:

```yaml
name: decosa
x-gpu: &gpu
  deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG:-0.1.0}
    build: docker/llm
    <<: *gpu
    ipc: host
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL:-nvidia/Qwen3.8-27B-NVFP4}", "--revision", "${LLM_REVISION:-482ca0f3832238542f8f5295dde86b5f22711d80}",
              "--served-model-name", "qwen3.8-27b", "--language-model-only",
              "--max-model-len", "${LLM_MAX_LEN:-65536}", "--gpu-memory-utilization", "${LLM_GPU_UTIL:-0.60}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3",
              "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}', "--seed", "0",
              "--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 5, start_period: 900s }
  asr:
    image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG:-0.1.0}
    build: docker/asr
    <<: *gpu
    ipc: host
    depends_on: { llm: { condition: service_healthy } }
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
              "--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
              "--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL:-0.25}",
              "--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 5, start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG:-0.1.0}
    build: { context: ., dockerfile: docker/api/Dockerfile }
    depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
    environment:
      DECOSA_ASR_WS: ws://asr:8000/v1/realtime
      DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE:-direct}
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP:-4}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_AUDIO_S: "3600"
      DECOSA_BUDGET_LLM_TOKENS: "200000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CORS_ORIGINS: ${DECOSA_CORS_ORIGINS:-}
    ports: ["${DECOSA_BIND:-127.0.0.1}:${DECOSA_PORT:-8445}:8445"]
    volumes: [decosa-data:/data]
volumes: { hf-cache: {}, decosa-data: {} }
```

For the two-card layout, give `asr` its own device: replace `<<: *gpu` in `asr` with the same block using `device_ids: ["1"]`. The shares must add up to under about 0.9 when both engines share one card.

## 4. Start and wait

```bash
docker compose pull || docker compose build   # images are "publishing soon"; build falls back to source
docker compose up -d
docker compose ps                              # llm healthy first (5-10 min on first start), then asr (1-3 min)
curl -s localhost:8445/healthz                 # expect "ok": true, "asr": true, "llm": true, "llm_route": "direct"
```

If `llm` stays unhealthy, read `docker compose logs llm`: usually out of memory (lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`) or NVFP4 on a pre-Blackwell GPU (switch to the FP8 checkpoint).

## 5. Smoke test: replay a sample walk-through

The replay route runs a canned 16 kHz recording through the real ASR and LLM pipeline and streams the same events as the live WebSocket, over SSE.

```bash
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"field"}' | jq -r .token)
curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"vertical":"field","script_id":"field-roof-inspection"}' | tee roof.sse | grep -o '"type": *"[a-z]*"' | sort | uniq -c
```

Pass when the stream has a `ready` event, `transcript` events, `lane` events for `checklist`, `issues`, `measurements` and `passed_checks` during the run, then one `lane` event each for `safety_sweep`, `report_check` and `report` after the audio ends, and a final `done`. Expect `receipt` events with `"status": "attested"` on the direct route; signed by the box's own key rather than a gateway. The other sample is `field-electrical-panel`. Lane events replace the lane's previous state; they do not append.

## 6. Export the report (JSON, markdown, PDF)

The `done` event carries `summary.report` (structured JSON: title, summary, overall_condition, issues with severity and priority, measurements with the stated spec and within_spec, passed_checks, code_references, site_details, checklist, recommendations, follow_up; every issue, measurement, passed check, code reference and site detail carries `quote` (the inspector's words, verbatim), `lines`, `t` (seconds into the walk-through) and `verified`) and `summary.report_markdown`.

```bash
grep '^data: ' roof.sse | sed 's/^data: //' | jq -c 'select(.type=="done") | .summary' > done.json
jq '.report' done.json > roof-report.json
jq -r '.report_markdown' done.json > roof-report.md
pandoc roof-report.md -o roof-report.pdf     # optional; needs pandoc and a PDF engine
jq -r '.issues[] | "\(.severity)\t\(.issue)\t\(.location)"' roof-report.json
```

## 7. Point the app at the local API

- Set the web app's API base to `http://localhost:8445` (for the Decosa web app, `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`).
- Live mic: open `ws://localhost:8445/ws/live?vertical=field&token=$TOKEN`, send 16 kHz mono PCM16 little-endian frames of about 100 ms, then the text frame `{"type":"stop"}`, and read the JSON events.
- If the page is served from another origin, add it to `DECOSA_CORS_ORIGINS` (a WebSocket close with code `4403` means the origin was refused).
- To reach it from phones on site, put a TLS reverse proxy on the LAN in front of port 8445 and set `DECOSA_TRUSTED_PROXIES`. Do not expose it to the internet.

When done, report the GPU tier you chose, the `/healthz` output, the lane ids seen in the smoke test, and the path of the exported report files.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Talk through an inspection, site visit or job; get the checklist, issues by severity, measurements and a finished report.
Who it's for
Teams in field and trades.
Where it runs
Hosted or self-host
Key numbers
  • 40/41 (98%); before 38/41 (93%) Expected issues in the report (grounded report + safety sweep + claim check) (synthetic, n = 41)
  • 40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6) Severity within the accepted range, of issues found (synthetic, n = 40)
  • 51/51; before 32/51 (63%) Stated facts kept (synthetic, n = 51)
  • 26.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Voxtral 4B · Qwen3.8-27B
Where
Hosted or self-host
Checks
Strong on text lanes
Output
Notes, reports and drafts · Structured data
Data
Personal data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

How does Decosa turn a spoken inspection into a report?

With Decosa field reports, the inspector narrates the walk-through into a phone or laptop. Speech is transcribed live, and about every 10 seconds a language model updates a checklist, an issue list graded critical, major, minor or info, measurements heard, and passed checks, each quoted from the transcript with its time. On stop it writes a JSON and Markdown report, runs a safety sweep and a claim check. The report is a draft for the inspector to review and sign.

How accurate are Decosa's AI inspection reports?

On Decosa's internal eval, field reports included 40/41 (98%) expected issues, rated 40/40 found issues within the accepted severity range with all 6 critical-only hazards rated critical, and kept 51/51 stated facts; the judge flagged 5/125 claims. The caveats matter: 8 synthetic TTS walk-throughs, no real site recordings, prompts tuned on these same scripts so there is no held-out set, and a Gemma 4 judge checked by hand only in part.

Does the inspection report app work offline on a job site?

Decosa field reports can be self-hosted: the whole stack runs on one GPU box on-site and keeps working without an internet connection once the model weights are downloaded. With DECOSA_LLM_ROUTE=direct, nothing leaves the box. A lite tier runs on a 32 GB Blackwell card plus a card of 16 GB or more for speech, with 16k context and 1 to 2 walk-throughs at a time.

Does it check measurements against building code?

No. Decosa field reports check each measurement against the spec the inspector states aloud; there is no built-in code database. The report is a drafting aid for the inspector to review and sign, and it does not replace a licensed inspection or a code-compliance determination. Self-hosting keeps site audio, transcripts and reports on your own hardware.

How long does the report take and what does it cost?

Decosa field reports write the final report after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load (hosted p50 of 26 s over 3 replays). That 75 s roof sample cost about $0.026 at the gateway list price, about 41k tokens and 28 model calls. Output is structured JSON and Markdown, with PDF via pandoc.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Field reports

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.