Eval: agent flight recorder (vertical 26)
Run on 25 Sep 2026 on our server: decosa-api branch the pre-release branch on 127.0.0.1:8486, gateway route (Qwen3.8-27B,
every decision receipted), signing with the live instance key. Script: scripts/flight_eval.py. Numbers:
docs/evals/flight-recorder-results.json (every tamper trial is listed there).
What was measured
- Verify rate on genuine runs. 19 sealed records from four sources:
- 3 hosted-demo runs (headless Chromium on the server, fictional shop; the Watch fixture);
- 8 held-out agent runs through the SDK over HTTP: 4 tasks x 2 repeats, 2 tasks on the public Sauce Labs demo site (saucedemo.com, public test credentials) and 2 on the fictional shop;
- 7 real Jev macos-harness runs from the owner's machine, imported with
jev_macos.pyin hash-only mode (no image left the Mac; these records are not published); - 1 OpenAI computer-use shaped loop (synthetic items; no third-party API was called).
- Tamper detection. 22 kinds of alteration, each applied to a fresh copy of every record it applies to, then
checked through
POST /flight/verify: 339 trials. - Overhead per step. 40 steps with a 1280x800 JPEG screenshot (81 KB mean), posted through the SDK over local HTTP, with images and in hash-only mode; plus the in-process append.
- Agent success (secondary). Whether the demo agent reached the expected outcome on the held-out tasks.
Results
| Result | |
|---|---|
| Genuine records that verify | 19/19 (all also pass the pinned-signer check) |
| Tampered copies caught, issuer key pinned | 339/339 |
| Tampered copies caught without key pinning | 312/339: the 27 misses are the two "rebuild the whole chain and re-sign with another key" attacks, which only pinning can catch (by design) |
| Step located, for alterations inside a step | 100% for edited observations, decisions (even with the text hash fixed), actions, results, guards, thumbnails, receipts, deleted and swapped steps |
| Record check time, 44 entries and 11 thumbnails | 5 ms (server) |
| Recording overhead per step, with a screenshot | 18.3 ms median (p95 19.8 ms) over local HTTP; 110 KB sent; 15.3 KB of record (3.3 KB entries + 11.8 KB thumbnail) |
| Recording overhead per step, hash-only | 8.4 ms median (p95 9.8 ms); 2 KB sent; 3.3 KB of record |
| In-process append (no HTTP) | 8.8 ms with a server-made thumbnail; 0.07 ms hash-only |
| Decision latency (Qwen3.8-27B via the gateway, 83 decisions) | median 7.8 s, 10-90% 6.7-12.8 s, while the shared model server had 20-30 requests running and 50-70 queued by other evals |
| Decisions with a gateway-signed receipt | 83/83 |
| Held-out agent runs reaching the expected outcome | 8/8 (the two "finish the order" runs were correctly stopped by the approval guard before "Finish") |
Per alteration (detected / trials; "no pin" = detected without key pinning):
| Alteration | Detected | No pin | Where it points |
|---|---|---|---|
| Observation text edited | 16/16 | 16 | the observation |
| Observed URL edited | 12/12 | 12 | the observation |
| Decision text edited | 19/19 | 19 | the decision |
| Decision text edited, text hash fixed | 19/19 | 19 | the decision (entry hash) |
| Action target edited, entry hash recomputed | 18/18 | 18 | the next entry (broken link) |
| Action value edited | 18/18 | 18 | the action |
| Result flipped | 18/18 | 18 | the result |
| Guard changed to approved | 8/8 | 8 | the guard |
| Guard entry deleted | 8/8 | 8 | the entry after the gap (the run end when the guard was last) |
| Middle step deleted | 18/18 | 18 | the entry after the gap |
| Two steps swapped | 18/18 | 18 | the first moved entry |
| Last step and run end cut off | 18/18 | 18 | count, head and root |
| Outcome in the signed statement changed | 19/19 | 19 | signature |
| Task text edited | 19/19 | 19 | the run header |
| One thumbnail pixel changed | 12/12 | 12 | every step that uses that thumbnail |
| Two thumbnails swapped | 12/12 | 12 | both steps |
| Unreferenced image added | 19/19 | 19 | images check |
| Receipts swapped between decisions | 11/11 | 11 | both decisions (receipt covers a different output) |
| Decision receipt stripped | 11/11 | 11 | the decision |
| Signature replaced | 19/19 | 19 | signature |
| Decision rewritten, chain rebuilt, re-signed with another key | 19/19 | 0 | signer pin only |
| Guard turned into an approval, chain rebuilt, re-signed with another key | 8/8 | 0 | signer pin only |
Honest limits
- The record proves what was reported, not what happened. A client-reported step is only as honest as the agent.
Steps observed by the server's own browser are marked
server-browser; decisions made through/decidecarry the gateway receipt. Imported runs (source: imported) are only tamper-evident from the time of import. - Whoever holds the signing key can rewrite history. Without key pinning, a record rebuilt and re-signed with a
different key verifies (0 of 27 caught). Pin the issuer's key (
GET /attest/signing-key); the site's viewer does this automatically. Keep the key away from the party the record must hold to account. - Deleting all images is allowed. A record shared without attachments still verifies; the hashes remain, so a stored screenshot can still be proven later.
- Agent success is secondary and small. 8/8 on four held-out tasks is not a benchmark. For a benchmark of the decision model on its own (MiniWoB++ and Mind2Web, with and without the guards, and a
vision variant), see
computer-use-bench.md. The decision prompt and the "done" rule were adjusted on the three hosted demo tasks (the first run said "done" early and the completion check caught it); the held-out tasks were not used for any change. The model reads the element table, not pixels (the served model is language-only), so canvas-heavy or iframe-heavy apps need another observation source. - Latency was measured under heavy shared load. Decision times will be lower on a quiet GPU; the recorder's own overhead (8-18 ms) does not depend on it.
- A key's rate limit (60 requests a minute) bounds a single agent at about 30 steps a minute with
/decide, or 60 with client-side decisions. Each step re-writes the run file, so very long runs with images (hundreds of steps, 4 MB of images) cost more per step; an append-only log is the fix if that matters.
Fixes after blind developer testing (28 Sep 2026)
Two agent builders tried the recorder cold on their own synthetic runs and found two bugs; both are fixed, with tests
(tests/test_flight.py, test_rejected_step_writes_nothing, test_decision_keeps_action_target_reason; the file passes
26/26 plus the auditor file 19/19):
- A rejected step no longer writes anything. A POST with a valid observation and an invalid decision returned 400 but left the observation in the signed chain. One POST is now all-or-nothing (the chain, step counters, decision counts and images are restored when any part is rejected).
- Decisions keep what the agent chose.
{action, target, reason, confidence}sent with a client-reported decision were dropped, so the record could not say what the agent decided or why. They are now fields of thedecideentry (other unknown fields go inextra), and are hashed as the decision's text when no raw model output is sent. The tamper numbers above are unchanged: the fixes touch what is written, not how entries are chained or verified.