26 · Software and AI ops · Compliance and trust · live
Agent flight recorder
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Genuine records that verify19/19syntheticn = 193 hosted-demo, 8 held-out SDK runs, 7 imported Jev runs, 1 synthetic OpenAI-shaped loop.
- Tampered copies caught, issuer key pinned339/339syntheticn = 33922 kinds of alteration.
- Tampered copies caught without key pinning312/339syntheticn = 339The 27 misses are chains rebuilt and re-signed with another key, which only pinning catches (by design).
- Held-out agent runs reaching the expected outcome8/8held outn = 84 tasks x 2 repeats; not a benchmark.
- MiniWoB++ success, production agent with guards25.9%held outn = 62595% CI 22.6-29.5%; 35.2% without the guards; no tuning for the bench.
- Mind2Web element accuracy / step success44.7% / 40.7%held outn = 300Raw model answers; 35% step success under the production value guard.
- Recording overhead per step with a screenshot (median)18.3 mssyntheticn = 408.4 ms hash-only.
Dataset
19 sealed flight records (hosted demo, held-out SDK runs on saucedemo.com and a fictional shop, imported Jev macOS runs) with 339 tamper trials; plus the demo agent on MiniWoB++ (125 tasks x 5 seeds) and a 300-step Mind2Web sample.
Caveats
- The record proves what was reported, not what happened: a client-reported step is only as honest as the agent.
- Without key pinning, a record rebuilt and re-signed with another key verifies (0 of 27 caught).
- Agent success of 8/8 on four held-out tasks is small; the decision prompt was adjusted on the three hosted demo tasks.
- On public benchmarks the agent is demo-grade: 25.9% on MiniWoB++, failing on canvas, drag, custom widgets and unquoted values.
- MiniWoB++ is tiny and synthetic and Mind2Web is offline; no live multi-page benchmark (WebArena) was run.
- Decision latency was measured under heavy shared load.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 5.3 s
- Receipts
- 8
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 25 Sep 2026: Fresh git clone of decosa-api (ba02fab), api image built from docker/api/Dockerfile (705 MB, no browser), compose from the assemble prompt with the llm service dropped and DECOSA_LLM_URL pointed at an already-running Qwen3.8-27B vLLM on the same box.
Verified on 2026-09-25: the image builds, the service starts, and the smoke test passes end to end against a local model server equivalent to the documented one; model-server startup itself was not re-verified. Key minting, the prompt's smoke script (verify, then fail at step 1 after a change), /decide (click on element 1, receipt status attested, 238 ms), 20-step timing (18.5 ms per step with a 480 KB screenshot, 3.9 ms hash-only) and the site's record viewer pointed at the box all worked. Worked around locally (fixed in the shared self-host pass): the compose health check calls curl, which the image does not have.
Rehearsal bundle: flight-recorder.zip (39 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted demo sessions are limited per network each hour (the current number is in GET /healthz); the console also takes an API key.
- A key has a per-minute request rate and a cap on open runs (GET /flight/info lists the limits). Close an abandoned run with DELETE /flight/runs/<id>, or see your runs with GET /flight/runs; the Python SDK waits out a 429.
- Hosted decision latency depends on the shared model server: under a second per decision when quiet, several seconds when busy.
- A restart of the hosted service stops a demo run in progress (shown as 'the demo run stopped').
- The record proves what was reported and that it was not changed after signing; it does not prove a website did what it showed.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDKAGPL-3.0-or-later (the SDK, the flight record format and its verifier are Apache-2.0)
- Decision model: picks the next action from a numbered element table (never coordinates, never free text)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Headless browser for the hosted demo and the Playwright adapterPlaywright 1.58 with Chromium headless shellApache-2.0 (Playwright); BSD-3-Clause (Chromium)
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
How well does the agent do on its own?
- MiniWoB++, 125 small web tasks x 5 seeds, production agent with guards: 25.9% (22.6-29.5%) (36.7% on tasks built from links, buttons and inputs; 4% on canvas, shape and colour tasks; 4% on drag, slider and keyboard tasks)
- Same, without the guards: 35.2% (31.6-39.0%) (The 9-point gap is almost all the value guard: it refused to type dates, sums and other values the task did not quote)
- Same weights reading the screenshot instead of the table (not served today): 55.8% (51.9-59.7%) (Pixels only, coordinate clicks, no guards; 0.75 s per decision on a private server)
- Mind2Web, 300 recorded steps on real websites: right element / right element and operation: 44.7% / 40.7% (Brackets 39.1-50.3% and 35.3-46.3%. Under the production value guard, 35% of steps would go through)
- Decision time, production agent: 0.59 s median (p90 7.1 s when the shared model server was busy)
Source: decosa-api docs/evals/computer-use-bench.md and computer-use-bench.json (every episode and model answer), 2026-09-25. MiniWoB++ (MIT) through BrowserGym (Apache-2.0); Mind2Web (CC BY 4.0) with the Multimodal-Mind2Web test subset (OpenRAIL). WebArena was not run: it needs six self-hosted sites.
Lite · record only, any CPU (3)
- Genuine records that verify: 19/19 (3 hosted demo runs, 8 SDK agent runs, 7 imported Jev harness runs, 1 computer-use loop)decosa-api docs/evals/flight-recorder.md, 2026-09-25
- Tampered copies caught (22 kinds of alteration): 339/339 with the issuer key pinned; 312/339 without (the 27 were rebuilt and re-signed with another key, which only pinning can catch)decosa-api docs/evals/flight-recorder-results.json
- Overhead per step: 18 ms and 15 KB of record with a thumbnail; 8 ms and 3.3 KB hash-onlydecosa-api docs/evals/flight-recorder.md (40-step benchmark, local HTTP)
Standard · receipted decisions, one 96 GB card (hosted demo) (5)
- Held-out agent tasks reaching the expected outcome (2 on saucedemo.com, 2 on the demo shop, 2 runs each; 2 expected a guard stop): 8/8decosa-api docs/evals/flight-recorder.md; the prompt was adjusted on the three demo tasks, not these
- Decisions with a gateway-signed receipt: 83/83decosa-api docs/evals/flight-recorder.md
- Guard stops on an order or payment button: 3/3 runs whose task asked to place or finish an order stopped before the clickdecosa-api docs/evals/flight-recorder.md
- MiniWoB++ success, the agent on its own (125 tasks x 5 seeds, BrowserGym task classes): 25.9% (95% CI 22.6-29.5%) with the guards; 35.2% (31.6-39.0%) without them; 36.7% on the form-and-button tasks, 4% on canvas, drag and slider tasksdecosa-api docs/evals/computer-use-bench.md, 2026-09-25; production prompt, no tuning
- Mind2Web, next-step accuracy on real websites (300 test steps, top-50 candidates): element 44.7% (39.1-50.3%), step success 40.7% (35.3-46.3%) before the guards; 35% after themdecosa-api docs/evals/computer-use-bench.md, 2026-09-25