Skip to content
decosa
LiveHostedSelf-hostMac

Prove what your agent did

A signed, step-by-step record of what your agent saw, decided and did, which anyone can re-check. Change one step later and verification names it.

Measured339/339Edited copies of a record caught, signing key pinned (synthetic set)
On production5.3 smedian on production (2026-09-25); slower when the service is busy
List price~$0.25 per 100 runsmeasured, at list price

Built on: Agent flight recorder, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the agent flight recorder API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for the decision model; the recorder and browser run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
flight-recorder

Use the hosted API

# Decosa agent flight recorder: record my agent with the hosted API

You are adding Decosa's flight recorder to the agent in this project. Each step the agent takes (what it saw, what it
decided, what it did, what happened) is posted to the recorder, chained, and sealed into a signed record anyone can
re-check. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`. What the recorder does and does not prove: `GET https://api.decosa.ai/flight/info`.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, enabled for `flight-recorder`. Keep it in
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`. A key allows 60 requests a minute.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "flight-recorder"}` returns `{"token", "expires_at",
   "budget"}`. a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); 429 with `Retry-After` over a limit.

## Privacy first
The hosted recorder keeps runs for 24 hours. If the agent's screens can show personal data, construct the SDK with
`send_images=False` (only screenshot hashes leave this machine; keep the screenshots yourself) and ask me whether to
self-host instead. Never put passwords in the task text; values typed into password fields are recorded as a hash.

## SDK (Python, standard library only)
`curl -fsSL https://api.decosa.ai/flight/sdk/decosa_flight.py -o decosa_flight.py` and, for Playwright agents,
`curl -fsSL https://api.decosa.ai/flight/sdk/playwright_agent.py -o playwright_agent.py`. Other adapters: `jev_macos.py`
(macos-harness), `openai_cu.py` (computer_call loops).

```python
import os
from decosa_flight import FlightRecorder
from playwright_agent import observe_page          # screenshot + numbered element table + page text

rec = FlightRecorder("https://api.decosa.ai", key=os.environ["DECOSA_API_KEY"])   # send_images=False for hash-only
rec.start(task, agent={"name": "my-agent", "harness": "playwright", "model": "my-model"})
for n in range(1, 20):
    obs = observe_page(page)
    raw = my_model(task, obs)                        # the agent's own model call, unchanged
    rec.step(observation={"url": obs["url"], "title": obs["title"], "screenshot": obs["screenshot"], "text": obs["text"]},
             decision={"model": "my-model", "output": raw}, step=n)
    ok, err = my_execute(page, raw)                  # unchanged
    rec.step(action={"type": "click", "target": {"name": "..."}, "description": "..."},
             result={"ok": ok, "error": err, "url_after": page.url}, step=n)
    if done: break
record = rec.seal("done", verified={"ok": my_check(page), "check": "what my check looked at"})
print(rec.verify(record)["summary"])                 # keep `record` (JSON) with the job it belongs to
```

To have the decision made by the receipted open model instead of your own: `rec.decide(observation=..., elements=[{"i",
"role", "name", "value"?, "options"?}], page_text=..., history=[...], step=n)` returns `{decision: {action, element,
value, reason}, guard}`. Never execute a decision whose `guard` is set: it was blocked or needs a person.
`playwright_agent.run_task(page, rec, task, verify=...)` runs that whole loop.

## Endpoints
- `POST /flight/runs` (token) `{task, agent?, policy?: {allowed_values?, approved?}}` → `{run_id, ...}`.
- `POST /flight/runs/{id}/steps` (token) `{step?, observation?, decision?, guard?, action?, result?, note?}`. Phases of a
  step go forward; a new step must be higher. Images: `screenshot_b64` (≤ 3 MB; the server keeps a 320 px thumbnail and
  the hash) or `screenshot_sha256` for hash-only. A decision made through this server's `/v1/chat/completions` can
  carry `receipt_id`; the signed receipt is then embedded, but only if it covers exactly `output`.
- `POST /flight/runs/{id}/decide` (token) → a receipted Qwen3.8-27B decision plus guards (see above).
- `POST /flight/runs/{id}/seal` (token) `{outcome: done|blocked|escalated|failed|stopped|max_steps, verified?}` →
  `{record, check}`. `GET /flight/runs/{id}/record` returns it again.
- `POST /flight/verify` `{record}` (no token) → `{ok, summary, checks, bad, steps, first_bad_step}`.

## Rules
- Record every step, including failed actions and guard stops. Do not drop steps to make a run look clean.
- The model's "done" is not success: pass `verified` from a check that reads the page or the system of record.
- Show users the summary and the first failing step when verification fails; never hide a failed check.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa agent flight recorder: run it yourself (containers)

You are setting up Decosa's flight recorder on this machine, so agent runs and their screenshots stay here. The
recorder (decosa-api) needs no GPU. Qwen3.8-27B is only needed if agents should get receipted decisions from
`/decide`; agents that bring their own model do not need it. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: any Linux x86_64 machine for the recorder; 1x RTX PRO 6000 96 GB for Qwen3.8-27B (NVFP4) if you want
`/decide`. Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/flight-recorder.zip (39 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py flight-recorder` (the api image carries the same bundle under /app/rehearsal/flight-recorder/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py flight-recorder --bundle flight-recorder.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "on the cart page the model chooses to click Checkout", "on the review page the Place order click is caught by the needs_approval guard", "the agent is stopped and the run escalated to a person"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Rules you must keep
- Bind every port to 127.0.0.1. Agents on other machines reach it over a VPN or an authenticated proxy, never directly.
- The signing key is created on first start in the data volume (`attest/ed25519.pem`, 0600). Back it up; never print
  it. If the people running the agents are the ones the record must hold to account, keep this box and its key out of
  their hands.
- Keep `DECOSA_FLIGHT_TTL_S` as short as the business allows; runs are deleted after it. Export sealed records to your
  own archive (they verify offline).

## Steps
1. Docker (and the NVIDIA Container Toolkit if you run the model): if `docker compose version` fails, install it from
   the official Docker instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `api` service; keep `llm` only if I want `/decide`.
3. In `.env`: `DECOSA_LLM_ROUTE=direct` (with `llm`), `DECOSA_FLIGHT_TTL_S=86400`, `DECOSA_FLIGHT_DEMO=0` (the demo
   browser is not needed here). The image needs Pillow for thumbnails (the `flight` extra).
4. Start: `docker compose pull && docker compose up -d`. Wait for `curl -fsS http://localhost:<PORT>/healthz`.
5. Mint a key for the agent: `POST /v1/keys` with the admin secret and `{"label": "agents", "verticals": ["flight-recorder"]}`.
6. Smoke test: `curl -fsSL localhost:<PORT>/flight/sdk/decosa_flight.py -o decosa_flight.py`, then in Python start a run,
   post one step with a small PNG screenshot and one without images (`screenshot_sha256`), and seal it. Expect a
   record whose `/flight/verify` says `ok: true`. Change one character of an entry's text and verify again: it must fail
   and name that step.
7. If `llm` runs: post `/flight/runs/{id}/decide` with a two-element table and check the `decide` entry has status
   `attested` (signed by this box's key: an attestation by me, not a gateway receipt).
8. Report back: health, the signing key id from `/attest/signing-key`, the smoke-test verify summaries, and the
   per-step time you saw.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/flight-recorder-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"flight-recorder"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py flight-recorder

Download the mock-data bundle (39 KB, 9 checks)expected.json

A browser agent is asked to buy a mug on Harbor Supply, a fictional demo shop. Two observations (screenshot, element table and page text) go to the receipted decision model: on the cart page it must choose an action, and on the order review page the Place order click must be escalated by the needs_approval guard instead of executed. The run is sealed; the record must verify, and a copy whose recorded click target was changed must fail at step 1.

What the rehearsal checks
  • on the cart page the model chooses to click Checkout
  • on the review page the Place order click is caught by the needs_approval guard
  • the agent is stopped and the run escalated to a person
  • the guard event is chained in the sealed record
  • the sealed record verifies
  • both screenshots are in the record and match their recorded hashes
  • a record whose recorded click target was changed no longer verifies
  • verification points at step 1, where the change was made
  • every model call has a signed receipt

Licence: Synthetic: Harbor Supply is a fictional demo shop shipped with decosa-api (AGPL-3.0-or-later); the screenshots were rendered from its pages. No real shop, orders or people.

Prompt for your coding agent

# Decosa agent flight recorder: run it yourself (containers)

You are setting up Decosa's flight recorder on this machine, so agent runs and their screenshots stay here. The
recorder (decosa-api) needs no GPU. Qwen3.8-27B is only needed if agents should get receipted decisions from
`/decide`; agents that bring their own model do not need it. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: any Linux x86_64 machine for the recorder; 1x RTX PRO 6000 96 GB for Qwen3.8-27B (NVFP4) if you want
`/decide`. Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/flight-recorder.zip (39 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py flight-recorder` (the api image carries the same bundle under /app/rehearsal/flight-recorder/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py flight-recorder --bundle flight-recorder.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "on the cart page the model chooses to click Checkout", "on the review page the Place order click is caught by the needs_approval guard", "the agent is stopped and the run escalated to a person"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Rules you must keep
- Bind every port to 127.0.0.1. Agents on other machines reach it over a VPN or an authenticated proxy, never directly.
- The signing key is created on first start in the data volume (`attest/ed25519.pem`, 0600). Back it up; never print
  it. If the people running the agents are the ones the record must hold to account, keep this box and its key out of
  their hands.
- Keep `DECOSA_FLIGHT_TTL_S` as short as the business allows; runs are deleted after it. Export sealed records to your
  own archive (they verify offline).

## Steps
1. Docker (and the NVIDIA Container Toolkit if you run the model): if `docker compose version` fails, install it from
   the official Docker instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `api` service; keep `llm` only if I want `/decide`.
3. In `.env`: `DECOSA_LLM_ROUTE=direct` (with `llm`), `DECOSA_FLIGHT_TTL_S=86400`, `DECOSA_FLIGHT_DEMO=0` (the demo
   browser is not needed here). The image needs Pillow for thumbnails (the `flight` extra).
4. Start: `docker compose pull && docker compose up -d`. Wait for `curl -fsS http://localhost:<PORT>/healthz`.
5. Mint a key for the agent: `POST /v1/keys` with the admin secret and `{"label": "agents", "verticals": ["flight-recorder"]}`.
6. Smoke test: `curl -fsSL localhost:<PORT>/flight/sdk/decosa_flight.py -o decosa_flight.py`, then in Python start a run,
   post one step with a small PNG screenshot and one without images (`screenshot_sha256`), and seal it. Expect a
   record whose `/flight/verify` says `ok: true`. Change one character of an entry's text and verify again: it must fail
   and name that step.
7. If `llm` runs: post `/flight/runs/{id}/decide` with a two-element table and check the `decide` entry has status
   `attested` (signed by this box's key: an attestation by me, not a gateway receipt).
8. Report back: health, the signing key id from `/attest/signing-key`, the smoke-test verify summaries, and the
   per-step time you saw.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/flight-recorder-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsAgent flight recorder on GeForce RTX 5090: use the Standard · receipted decisions, one 96 GB card (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · receipted decisions, one 96 GB card (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Recorder: decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK. CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Decision model: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'.
  • Headless browser for the hosted demo and the...: Playwright 1.58 with Chromium headless shell. CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Agent flight recorder, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Agent flight recorder on my hardware

Fetch https://decosa.ai/prompts/flight-recorder-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=flight-recorder)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · receipted decisions, one 96 GB card (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Recorder: decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK, CPU
- Decision model: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Headless browser for the hosted demo and the...: Playwright 1.58 with Chromium headless shell, CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/flight-recorder-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --browser

Mac prompt for your coding agent

# Decosa Agent flight recorder: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Agent flight recorder on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/flight-recorder.zip (39 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py flight-recorder` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "on the cart page the model chooses to click Checkout", "on the review page the Place order click is caught by the needs_approval guard", "the agent is stopped and the run escalated to a person"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Decision model: picks the next action from a numbered element table (never coordinates, never free text) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Headless browser for the hosted demo and the Playwright adapter | CPU | Playwright's macOS arm64 Chromium | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --browser`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py flight-recorder`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --browser --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 5.3 s · ~$0.003 per run · 8 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh git clone of decosa-api (ba02fab), api image built from docker/api/Dockerfile (705 MB, no browser), compose from the assemble prompt with the llm service dropped and DECOSA_LLM_URL pointed at an already-running Qwen3.8-27B vLLM on the same box.

Measured cost to run: about $0.25 per 100 runs (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the image builds, the service starts, and the smoke test passes end to end against a local model server equivalent to the documented one; model-server startup itself was not re-verified. Key minting, the prompt's smoke script (verify, then fail at step 1 after a change), /decide (click on element 1, receipt status attested, 238 ms), 20-step timing (18.5 ms per step with a 480 KB screenshot, 3.9 ms hash-only) and the site's record viewer pointed at the box all worked. Worked around locally (fixed in the shared self-host pass): the compose health check calls curl, which the image does not have.

Known limits (5)
  • Hosted demo sessions are limited per network each hour (the current number is in GET /healthz); the console also takes an API key.
  • A key has a per-minute request rate and a cap on open runs (GET /flight/info lists the limits). Close an abandoned run with DELETE /flight/runs/<id>, or see your runs with GET /flight/runs; the Python SDK waits out a 429.
  • Hosted decision latency depends on the shared model server: under a second per decision when quiet, several seconds when busy.
  • A restart of the hosted service stops a demo run in progress (shown as 'the demo run stopped').
  • The record proves what was reported and that it was not changed after signing; it does not prove a website did what it showed.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the agent flight recorder API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for the decision model; the recorder and browser run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

A signed, step-by-step record of what a browser or computer-use agent saw, decided and did, that anyone can re-check.

Your agent posts each step to the recorder: the page or screen it saw (a screenshot hash and thumbnail, and the element table or accessibility tree the model was shown), the model's decision, the action and the result. Every entry is appended to a hash chain; when the run ends it is sealed with an Ed25519 signature. Decisions can run on Qwen3.8-27B through our gateway, and then the signed receipt for that exact output sits inside the step. Guards run in code before anything is executed: a typed value must come from the task, and an order, payment or deletion waits for a person. A viewer replays the run step by step, and verification in the browser names the step where a screenshot, a decision or a guard was changed. It is for teams running agents on back-office work who need evidence after an incident, in a customer dispute or for an audit.

Deployment
Hosted or self-host
Regulatory
Checked 2026-09-25. EU AI Act: Article 12 requires high-risk AI systems to be able to record events automatically (logs) over their lifetime, and Articles 19 and 26(6) require providers and deployers to keep those logs for at least six months unless other law says otherwise. That duty applies only to high-risk systems (Annex I products, and Annex III uses such as hiring, credit scoring or access to essential services); after the AI Omnibus (Council approval 29 June 2026) Annex III duties apply from 2 December 2027 and Annex I from 2 August 2028. Most browser agents doing back-office work are not high-risk systems, and for them this record is useful evidence, not a legal requirement. Article 12 does not require signatures, hash chains or screenshots, and the harmonised logging standards (prEN 18229-1, ISO/IEC DIS 24970) are still drafts, so no product can claim conformity with them yet; this one does not. Screenshots of back-office screens usually contain personal data: under the GDPR, minimisation and storage limits apply, so use hash-only mode (images stay with you) or self-host. The hosted demo keeps runs for 24 hours and only visits a fictional shop. In a dispute a signature shows the record was not changed after signing and which key signed it; it does not show that what an agent reported about itself was true, and its weight as evidence is for the court or arbitrator. Model licence: Apache-2.0 (Qwen3.8-27B). Not legal advice.
Architecture
Text description

Agents (a Playwright agent, the Jev macos-harness, an OpenAI-style computer-use loop, or the hosted demo agent) send each step to the recorder over HTTP or the SDK. The recorder chains five kinds of entry per step: what the agent saw, what it decided, guard events, what it did and the result. Decisions can run on Qwen3.8-27B through our gateway, which returns a signed receipt that is embedded in the decision entry. Guards in code stop typed values that are not in the task and orders, payments or deletions that need a person. The sealed record has a Merkle root, signed checkpoints and an Ed25519 signature, with thumbnails attached by hash. Anyone can verify it in a browser or at POST /flight/verify, and a change to a pixel, a word or a step fails and names the step. The recorder, screenshots and signing key stay on your box when self-hosted.

Architecture

At a glance

Data retention
Hosted: runs and screenshots are kept 24 hours (DECOSA_FLIGHT_TTL_S), then deleted. Self-host: you set the TTL; sealed records verify offline, so archive them yourself.
What leaves the box
Hosted: whatever the agent posts (screenshots, page text, decisions), unless the SDK runs with send_images=False, which sends only screenshot hashes. Self-host: nothing, unless you point DECOSA_LLM_ROUTE at a gateway.
Inputs and limits
Screenshots as JPEG, PNG or WebP up to 3 MB each (4 MB of images per run; the server keeps a 320 px thumbnail and the hash), page text up to 16,000 characters per step, up to 300 steps per run.
Logs
The service log carries method, path, status and timing only; task text and screenshots are not logged (checked on the self-host box).
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    record only, any CPU

    Your agent and your model; the recorder chains, seals and verifies. Decisions from another model are recorded with their hash but carry no receipt.

    Models
    • decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK
    Hardware
    Any Linux or macOS machine with Python 3.11+
    Quality evidence
    • Genuine records that verify19/19 (3 hosted demo runs, 8 SDK agent runs, 7 imported Jev harness runs, 1 computer-use loop)decosa-api docs/evals/flight-recorder.md, 2026-09-25
    • Tampered copies caught (22 kinds of alteration)339/339 with the issuer key pinned; 312/339 without (the 27 were rebuilt and re-signed with another key, which only pinning can catch)decosa-api docs/evals/flight-recorder-results.json
    • Overhead per step18 ms and 15 KB of record with a thumbnail; 8 ms and 3.3 KB hash-onlydecosa-api docs/evals/flight-recorder.md (40-step benchmark, local HTTP)
    Latency
    measured: milliseconds per step on the recorder; the rest is your agent
    Verification
    Proof: partialSelf-host onlyThe record is attested by the box's own key; decisions from a model outside this server are not receipted.
  • In the hosted demo

    Standard

    receipted decisions, one 96 GB card (hosted demo)

    Qwen3.8-27B makes each decision through /decide, with a gateway receipt inside the step and guards before anything runs.

    Models
    • decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK
    • Qwen3.8-27B (NVIDIA NVFP4)
    • Playwright 1.58 with Chromium headless shell
    Hardware
    1× RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Held-out agent tasks reaching the expected outcome (2 on saucedemo.com, 2 on the demo shop, 2 runs each; 2 expected a guard stop)8/8decosa-api docs/evals/flight-recorder.md; the prompt was adjusted on the three demo tasks, not these
    • Decisions with a gateway-signed receipt83/83decosa-api docs/evals/flight-recorder.md
    • Guard stops on an order or payment button3/3 runs whose task asked to place or finish an order stopped before the clickdecosa-api docs/evals/flight-recorder.md
    • MiniWoB++ success, the agent on its own (125 tasks x 5 seeds, BrowserGym task classes)25.9% (95% CI 22.6-29.5%) with the guards; 35.2% (31.6-39.0%) without them; 36.7% on the form-and-button tasks, 4% on canvas, drag and slider tasksdecosa-api docs/evals/computer-use-bench.md, 2026-09-25; production prompt, no tuning
    • Mind2Web, next-step accuracy on real websites (300 test steps, top-50 candidates)element 44.7% (39.1-50.3%), step success 40.7% (35.3-46.3%) before the guards; 35% after themdecosa-api docs/evals/computer-use-bench.md, 2026-09-25
    Latency
    measured: seconds per decision while the shared GPU was saturated; about a minute for a short run
    Verification
    Proof: strongEvery decision made through /decide is gateway-receipted and bound to its output hash in the chain.

Also runs on

  • Typed decisions on Apple silicon (Jev-compatible)DiffusionGemma 26B-A4B (MLX, OptiQ 4-bit)self-host onlyDiffusionGemma as a local decision server for the macOS harness: fast single choices, weak at ordered sequences. Not a hosted model, so no receipts. Hardware: Apple silicon with 32 GB or more.

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK
0 GBProof: partial
Decision model: picks the next action from a numbered element table (never coordinates, never free text)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57.6 GBProof: strongIn the hosted demo
Headless browser for the hosted demo and the Playwright adapterPlaywright 1.58 with Chromium headless shell
0 GBNo proof yetIn the hosted demo
Typed decisions on Apple silicon (Jev-compatible local decision server), for the macos-harness adapterDiffusionGemma 26B-A4B (MLX, OptiQ 4-bit)mlx-community/diffusiongemma-26B-A4B-it-OptiQ-4bit on Hugging Face (opens in a new tab)
26B (4B active) · 18.5 GBNo proof yet

Measured

How well does the agent do on its own?

The recorder's numbers above measure the record. This measures the demo agent: Qwen3.8-27B choosing actions from the numbered element table, on two public benchmarks, with the production prompt and no tuning. 25 Sep 2026, 95% intervals in brackets.

MiniWoB++, 125 small web tasks x 5 seeds, production agent with guards
25.9% (22.6-29.5%)36.7% on tasks built from links, buttons and inputs; 4% on canvas, shape and colour tasks; 4% on drag, slider and keyboard tasks
Same, without the guards
35.2% (31.6-39.0%)The 9-point gap is almost all the value guard: it refused to type dates, sums and other values the task did not quote
Same weights reading the screenshot instead of the table (not served today)
55.8% (51.9-59.7%)Pixels only, coordinate clicks, no guards; 0.75 s per decision on a private server
Mind2Web, 300 recorded steps on real websites: right element / right element and operation
44.7% / 40.7%Brackets 39.1-50.3% and 35.3-46.3%. Under the production value guard, 35% of steps would go through
Decision time, production agent
0.59 s medianp90 7.1 s when the shared model server was busy

What the element table means

The model never sees the page. It gets a numbered list of the links, buttons, inputs and selects it can act on, and picks one. That makes each choice checkable and recordable, and it is why it does well on ordinary forms. It is also a crutch: anything not in the list does not exist for the agent. In 20% of MiniWoB episodes the list was empty at the first step, because the clickable things were plain spans and divs.

Where it fails

Canvas and drawn shapes, colours, drag and sliders, custom widgets such as date pickers, and content inside iframes, which the extractor does not enter. Long, exploratory sequences fail too: an eight-step flight booking scored 0 of 5 with every setup, and paging through tabs to find a link often loops until the step limit. In most of these cases it says it is blocked rather than guessing.

Why the guards matter

Without them, the model said it was done in 82 of 625 runs (13%) when the task had not finished, and it typed values nobody gave it. With them, a run is only marked successful when the page agrees, and anything typed comes from the task or the caller's allowed values. That costs capability, and the table shows how much. Pass the values a task needs as allowed_values rather than turning the guard off.

Verdict

Demo-grade on its own. It is usable for narrow, form-shaped jobs on sites with real links and inputs, when the values are supplied and a completion check is written for the task, as in the hosted demo and the 8/8 held-out runs. It is not a general web agent. The vision result shows where the gain is: a hybrid that uses the table when it has the target and the screenshot when it does not, with the same guards.

Source: decosa-api docs/evals/computer-use-bench.md and computer-use-bench.json (every episode and model answer), 2026-09-25. MiniWoB++ (MIT) through BrowserGym (Apache-2.0); Mind2Web (CC BY 4.0) with the Multimodal-Mind2Web test subset (OpenRAIL). WebArena was not run: it needs six self-hosted sites.

Around the models

Tools, services and hardware

Tools

  • decosa_flight.py (SDK) and adapters: playwright_agent.py, jev_macos.py, openai_cu.pyApache-2.0

    Served at GET /flight/sdk/<file>. Record any agent over HTTP with a dk_ key; hash-only mode sends no images. The Jev adapter imports a macos-harness run directory or hooks its loop; the OpenAI adapter wraps computer_call and computer_call_output.

  • Harbor Supply (fictional demo shop)Apache-2.0

    Six static pages served to the demo browser by request interception. Nothing is sold; the completion checks read its session state.

  • Held-out eval tasks only (public test credentials printed on its login page). The hosted demo never visits it.

  • scripts/flight_eval.pyApache-2.0

    Held-out agent runs, overhead benchmark and the tamper set (decosa-api).

Services

  • decosa-api (flight routes):8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    POST /flight/runs, /flight/runs/{id}/steps, /decide, /seal; POST /flight/verify; POST /flight/demo (SSE); GET /flight/info, /flight/tasks, /flight/sdk/<file>.

  • Decision model (vLLM, or our gateway):8114
    vllm/vllm-openai:v0.29.0

    Qwen3.8-27B for /decide. Not needed when your agent brings its own model.

Hardware

  • Any CPU, no GPU Fits

    Recording, sealing and verification: 18 ms per step with a screenshot, 8 ms hash-only (local HTTP, measured).

  • 1× RTX PRO 6000 96 GB Fits

    Qwen3.8-27B NVFP4 for receipted decisions; measured on our server.

  • Apple silicon, 32 GB or more Fits

    DiffusionGemma MLX 4-bit for the Jev macos-harness: about 18.5 GB peak per request (measured on an M3 Ultra). Its decisions are not receipted.

Latency per lane

  • record one step (observation with a 1280x800 screenshot, decision, action, result), local HTTP18 ms

    Measuredmeasured on our server 2026-09-25: median of 40, p95 20 ms; 8 ms median in hash-only mode

  • one receipted decision (Qwen3.8-27B through the gateway)7.8 s

    Measuredmeasured on our server 2026-09-25: median of 83 decisions, 10-90% 6.7-12.8 s, while the shared model server had 20-30 requests running and 50-70 queued

  • hosted demo run, 8 steps (cheapest tool to review)80.6 s

    Measuredmeasured on our server 2026-09-25 (run fr-f2fbde7758536e72), same GPU load

  • verify a 44-entry record with 11 thumbnails5 ms

    Measuredmeasured on our server 2026-09-25 (server-side); the browser check runs the same algorithm with WebCrypto

Notes

  • The record proves what was reported to the recorder, in what order and when, which model made each receipted decision, and that nothing changed after signing. It does not prove that a website really did what it displayed, or that a step an agent reported about itself is true. Steps observed by the server's own browser are marked "seen by our browser".
  • The model's "done" never counts as success. A completion check reads the page itself; when it fails, a note goes on the record and the agent continues. The eval's agent runs reached the expected outcome in 8 of 8 held-out runs.
  • Someone who holds the signing key can rebuild and re-sign a whole record; someone who does not can only produce a record signed by a different key. Pin the issuer's key (GET /attest/signing-key) and verification catches that too. Keep the key off the machine the agent runs on if the agent's operator is the party you need to hold to account.
  • Hash-only mode keeps screenshots on your side: the record holds their sha256, so you can later prove a stored screenshot is the one the agent saw without ever sending it.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

flight-recorder/assemble-prompt.md151 lines
# Assemble the Decosa agent flight recorder on this machine

You are setting up a flight recorder for browser and computer-use agents: every step an agent takes (what it saw, what
it decided, what it did, what happened) is chained into a signed record that anyone can re-verify, and a change to one
screenshot, decision or step fails verification at that step. Optionally, agents get their decisions from Qwen3.8-27B
with a receipt per decision. Work step by step, show me each command before you run anything with `sudo`, and stop to
ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/flight-recorder.zip (39 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py flight-recorder` (the api image carries the same bundle under /app/rehearsal/flight-recorder/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py flight-recorder --bundle flight-recorder.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "on the cart page the model chooses to click Checkout", "on the review page the Place order click is caught by the needs_approval guard", "the agent is stopped and the run escalated to a person"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Recorder: decosa-api (AGPL-3.0-or-later), CPU only. Thumbnails: Pillow (MIT-CMU). Demo browser, optional: Playwright
  (Apache-2.0) with Chromium (BSD-3-Clause). Decision model, optional: Qwen3.8-27B (Apache-2.0).
- Screenshots of back-office screens hold personal data. Everything stays on this machine: bind every port to
  127.0.0.1, keep the data directory 0700, and prefer hash-only mode in the SDK so agents send only screenshot hashes.
- The signing key is the trust anchor. It is created on first start; back it up, never print it, and keep it away from
  the people whose agents are being recorded if the record must hold them to account.
- Be honest about what a record proves: what was reported, in what order and when, and that nothing changed after
  signing. Not that a website did what it displayed, and not that an agent's self-reported step is true.

## 1. Check the machine
1. Recorder only: any Linux x86_64 or macOS machine with Docker, or Python 3.11+. No GPU.
2. With receipted decisions: `nvidia-smi` shows one GPU with at least 32 GB (an RTX PRO 6000 96 GB runs Qwen3.8-27B
   NVFP4 with room to spare; an RTX 5090 32 GB fits it with a shorter context). Driver 570 or newer. On a card without
   NVFP4 use `Qwen/Qwen3.8-27B-FP8`.
3. `docker --version` and `docker compose version`. If Docker or (for the model) the NVIDIA container toolkit is missing,
   install them from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
4. Disk: 2 GB for the recorder image (about 400 MB more with the demo browser), about 25 GB more for the model, plus
   room for records (15 KB per step with thumbnails, about 3 KB hash-only).

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/flight/`, and `docker build -f docker/api/Dockerfile -t decosa-api:local .`. Add
  `--build-arg WITH_BROWSER=1` only if you want the hosted-style browser demo on this box.
- Without Docker: `python -m venv .venv && .venv/bin/pip install ".[flight]"` in the checkout, then
  `.venv/bin/python -m decosa_api`.
- For decisions: `vllm/vllm-openai:v0.29.0` and weights `nvidia/Qwen3.8-27B-NVFP4`.

## 3. docker-compose.yml
Write this in `~/decosa/flight/`. Drop the `llm` service and `depends_on` if agents bring their own model.

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--language-model-only", "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_FLIGHT_TTL_S: "86400"
      DECOSA_FLIGHT_DEMO: "0"
      DECOSA_ADMIN_SECRET: ${DECOSA_ADMIN_SECRET}
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/flight/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.
If you already run an
OpenAI-compatible Qwen3.8-27B server, drop `llm` and point `DECOSA_LLM_URL` at it; `DECOSA_LLM_MODEL` must be the
name that server serves (check `GET /v1/models`).

Put `DECOSA_ADMIN_SECRET=<a long random string>` in `~/decosa/flight/.env` (mode 0600). On the first start the api
service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Every decision made through `/decide` on
the direct route gets a receipt signed with that key (status `attested`): an attestation by me, the operator, not a
proof of computation.

## 4. Smoke test
1. `curl -s localhost:8445/flight/info | head -c 800` lists the entry kinds, limits, guards and the decision model.
2. Mint a key: `curl -s -XPOST localhost:8445/v1/keys -H "authorization: Bearer $DECOSA_ADMIN_SECRET" -H 'content-type: application/json' -d '{"label":"agents","verticals":["flight-recorder"]}'`.
   Keep the `dk_` key in the agent's environment as `DECOSA_API_KEY`.
3. `curl -s localhost:8445/flight/sdk/decosa_flight.py -o decosa_flight.py`, then run:
   ```python
   import base64, os, zlib, struct
   from decosa_flight import FlightRecorder
   def png(r, g, b):   # a 2x2 PNG, no dependencies
       raw = b"".join(b"\x00" + bytes([r, g, b]) * 2 for _ in range(2))
       c = lambda t, d: struct.pack(">I", len(d)) + t + d + struct.pack(">I", zlib.crc32(t + d) & 0xffffffff)
       return b"\x89PNG\r\n\x1a\n" + c(b"IHDR", struct.pack(">IIBBBBB", 2, 2, 8, 2, 0, 0, 0)) + c(b"IDAT", zlib.compress(raw)) + c(b"IEND", b"")
   rec = FlightRecorder("http://127.0.0.1:8445", key=os.environ["DECOSA_API_KEY"])
   rec.start("Smoke test: open the settings page.", agent={"name": "smoke"})
   rec.step(observation={"url": "https://app.example/", "screenshot": png(200, 30, 30), "text": "[1] link \"Settings\""},
            decision={"model": "none", "output": '{"action":"click","element":1}'},
            action={"type": "click", "target": {"index": 1, "name": "Settings"}}, result={"ok": True})
   rec.send_images = False
   rec.step(observation={"url": "https://app.example/settings", "screenshot": png(30, 200, 30)}, action={"type": "done"})
   record = rec.seal("done")
   print(rec.verify(record)["summary"])
   record["entries"][3]["target"]["name"] = "Delete account"
   print(rec.verify(record)["summary"])
   ```
   Expect `Verified: ...` and then `Verification failed at step 1 action ...`. The second step's screenshot is only a
   hash on the record.
4. With `llm`: post `/flight/runs/{id}/decide` with `{"observation": {"url": "https://app.example/"}, "elements":
   [{"i": 1, "role": "link", "name": "Settings"}], "page_text": "Home"}` on a run whose task says "open the settings
   page". Expect a `click` on element 1 and a `decide` entry with status `attested`. Tell me the latency you see.
5. Time the recorder: 20 steps with a real 1280x800 screenshot each. On our box the median was 18 ms per step with
   images and 8 ms hash-only. Tell me what you measure.

## 5. Point the agents at it
- Python agents: the SDK above (`FlightRecorder(API, key=..., send_images=False)`); Playwright agents also fetch
  `/flight/sdk/playwright_agent.py` (`observe_page(page)` gives the screenshot and element table in one call).
- The Jev macos-harness: fetch `/flight/sdk/jev_macos.py`; either hook the loop (`JevHook`) or import a finished run
  directory with `python jev_macos.py runs/<dir> --api http://127.0.0.1:8445 --key $DECOSA_API_KEY --hash-only`.
- OpenAI computer-use style loops: fetch `/flight/sdk/openai_cu.py` and wrap `computer_call` / `computer_call_output`.
- The site: set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in `.env.local` to use the viewer against this box.
  Contract: `API_CONTRACT.md`, section "Agent flight recorder".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that records screens with
personal data. If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my
explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

2 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A signed, step-by-step record of what a browser or computer-use agent saw, decided and did, that anyone can re-check.
Who it's for
Teams in software and ai ops and compliance and trust.
Where it runs
Hosted or self-host
Key numbers
  • 8/8 Held-out agent runs reaching the expected outcome (held out, n = 8)
  • 25.9% MiniWoB++ success, production agent with guards (held out, n = 625)
  • 44.7% / 40.7% Mind2Web element accuracy / step success (held out, n = 300)
  • 5.3 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B for decisions; any agent through the SDK
Where
Hosted or self-host
Checks
Receipted decisions; signed, hash-chained run record
Output
Signed record or verdict
Data
Personal data · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What goes into the AI agent audit trail?

The agent flight recorder takes each step your browser or computer-use agent posts: a screenshot hash and thumbnail, the element table or accessibility tree the model was shown, the model's decision, the action and the result. Every entry is appended to a hash chain and the run is sealed with an Ed25519 signature. If the decision ran on Qwen3.8-27B through our gateway, that output's signed receipt sits inside the step.

How does the agent flight recorder show that a record was changed?

A viewer replays the sealed run step by step, and verification in the browser names the step where a screenshot, a decision or a guard was changed. In the eval, 19/19 genuine records verified and 339/339 tampered copies were caught with the issuer key pinned. Without key pinning, 312/339 were caught: a chain rebuilt and re-signed with another key verifies, which only pinning catches.

Does the agent flight recorder meet EU AI Act Article 12?

No. EU AI Act Article 12 requires automatic event logs only for high-risk AI systems, and after the AI Omnibus Annex III duties apply from 2 December 2027. Most back-office browser agents are not high-risk. Article 12 does not require signatures, hash chains or screenshots, and the harmonised logging standards are still drafts, so the agent flight recorder claims no conformity with them. For most teams it is useful evidence, not a legal requirement.

What does an agent flight recorder record not prove?

The agent flight recorder proves what was reported to it, when, and that the record was not changed after signing. It does not prove that a website did what it showed, and a client-reported step is only as honest as the agent. In a dispute, a signature shows which key signed the record; its weight as evidence is for the court or arbitrator.

How well does the flight recorder's demo agent perform on its own?

The flight recorder's demo agent is demo-grade. With the production guards it scored 25.9% on MiniWoB++ (95% CI 22.6-29.5%) and 44.7% element accuracy on a 300-step Mind2Web sample, failing on canvas, drag, custom widgets and iframes. It reached the expected outcome in 8/8 held-out runs on four narrow tasks. The recorder is meant for recording your own agent, which costs nothing per step.

Where do screenshots go when I use the agent flight recorder?

Hosted, the agent flight recorder keeps runs and screenshots for 24 hours, then deletes them. With send_images=False the SDK sends only screenshot hashes, so images stay with you. Self-hosted, nothing leaves the box unless you point it at a gateway, and sealed records verify offline so you archive them yourself. Screenshots of back-office screens usually contain personal data, so GDPR minimisation applies.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Agent flight recorder

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.