Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: agent flight recorder (vertical 26)

Run on 25 Sep 2026 on our server: decosa-api branch the pre-release branch on 127.0.0.1:8486, gateway route (Qwen3.8-27B, every decision receipted), signing with the live instance key. Script: scripts/flight_eval.py. Numbers: docs/evals/flight-recorder-results.json (every tamper trial is listed there).

What was measured

  1. Verify rate on genuine runs. 19 sealed records from four sources:
    • 3 hosted-demo runs (headless Chromium on the server, fictional shop; the Watch fixture);
    • 8 held-out agent runs through the SDK over HTTP: 4 tasks x 2 repeats, 2 tasks on the public Sauce Labs demo site (saucedemo.com, public test credentials) and 2 on the fictional shop;
    • 7 real Jev macos-harness runs from the owner's machine, imported with jev_macos.py in hash-only mode (no image left the Mac; these records are not published);
    • 1 OpenAI computer-use shaped loop (synthetic items; no third-party API was called).
  2. Tamper detection. 22 kinds of alteration, each applied to a fresh copy of every record it applies to, then checked through POST /flight/verify: 339 trials.
  3. Overhead per step. 40 steps with a 1280x800 JPEG screenshot (81 KB mean), posted through the SDK over local HTTP, with images and in hash-only mode; plus the in-process append.
  4. Agent success (secondary). Whether the demo agent reached the expected outcome on the held-out tasks.

Results

Result
Genuine records that verify 19/19 (all also pass the pinned-signer check)
Tampered copies caught, issuer key pinned 339/339
Tampered copies caught without key pinning 312/339: the 27 misses are the two "rebuild the whole chain and re-sign with another key" attacks, which only pinning can catch (by design)
Step located, for alterations inside a step 100% for edited observations, decisions (even with the text hash fixed), actions, results, guards, thumbnails, receipts, deleted and swapped steps
Record check time, 44 entries and 11 thumbnails 5 ms (server)
Recording overhead per step, with a screenshot 18.3 ms median (p95 19.8 ms) over local HTTP; 110 KB sent; 15.3 KB of record (3.3 KB entries + 11.8 KB thumbnail)
Recording overhead per step, hash-only 8.4 ms median (p95 9.8 ms); 2 KB sent; 3.3 KB of record
In-process append (no HTTP) 8.8 ms with a server-made thumbnail; 0.07 ms hash-only
Decision latency (Qwen3.8-27B via the gateway, 83 decisions) median 7.8 s, 10-90% 6.7-12.8 s, while the shared model server had 20-30 requests running and 50-70 queued by other evals
Decisions with a gateway-signed receipt 83/83
Held-out agent runs reaching the expected outcome 8/8 (the two "finish the order" runs were correctly stopped by the approval guard before "Finish")

Per alteration (detected / trials; "no pin" = detected without key pinning):

Alteration Detected No pin Where it points
Observation text edited 16/16 16 the observation
Observed URL edited 12/12 12 the observation
Decision text edited 19/19 19 the decision
Decision text edited, text hash fixed 19/19 19 the decision (entry hash)
Action target edited, entry hash recomputed 18/18 18 the next entry (broken link)
Action value edited 18/18 18 the action
Result flipped 18/18 18 the result
Guard changed to approved 8/8 8 the guard
Guard entry deleted 8/8 8 the entry after the gap (the run end when the guard was last)
Middle step deleted 18/18 18 the entry after the gap
Two steps swapped 18/18 18 the first moved entry
Last step and run end cut off 18/18 18 count, head and root
Outcome in the signed statement changed 19/19 19 signature
Task text edited 19/19 19 the run header
One thumbnail pixel changed 12/12 12 every step that uses that thumbnail
Two thumbnails swapped 12/12 12 both steps
Unreferenced image added 19/19 19 images check
Receipts swapped between decisions 11/11 11 both decisions (receipt covers a different output)
Decision receipt stripped 11/11 11 the decision
Signature replaced 19/19 19 signature
Decision rewritten, chain rebuilt, re-signed with another key 19/19 0 signer pin only
Guard turned into an approval, chain rebuilt, re-signed with another key 8/8 0 signer pin only

Honest limits

  • The record proves what was reported, not what happened. A client-reported step is only as honest as the agent. Steps observed by the server's own browser are marked server-browser; decisions made through /decide carry the gateway receipt. Imported runs (source: imported) are only tamper-evident from the time of import.
  • Whoever holds the signing key can rewrite history. Without key pinning, a record rebuilt and re-signed with a different key verifies (0 of 27 caught). Pin the issuer's key (GET /attest/signing-key); the site's viewer does this automatically. Keep the key away from the party the record must hold to account.
  • Deleting all images is allowed. A record shared without attachments still verifies; the hashes remain, so a stored screenshot can still be proven later.
  • Agent success is secondary and small. 8/8 on four held-out tasks is not a benchmark. For a benchmark of the decision model on its own (MiniWoB++ and Mind2Web, with and without the guards, and a vision variant), see computer-use-bench.md. The decision prompt and the "done" rule were adjusted on the three hosted demo tasks (the first run said "done" early and the completion check caught it); the held-out tasks were not used for any change. The model reads the element table, not pixels (the served model is language-only), so canvas-heavy or iframe-heavy apps need another observation source.
  • Latency was measured under heavy shared load. Decision times will be lower on a quiet GPU; the recorder's own overhead (8-18 ms) does not depend on it.
  • A key's rate limit (60 requests a minute) bounds a single agent at about 30 steps a minute with /decide, or 60 with client-side decisions. Each step re-writes the run file, so very long runs with images (hundreds of steps, 4 MB of images) cost more per step; an append-only log is the fix if that matters.

Fixes after blind developer testing (28 Sep 2026)

Two agent builders tried the recorder cold on their own synthetic runs and found two bugs; both are fixed, with tests (tests/test_flight.py, test_rejected_step_writes_nothing, test_decision_keeps_action_target_reason; the file passes 26/26 plus the auditor file 19/19):

  • A rejected step no longer writes anything. A POST with a valid observation and an invalid decision returned 400 but left the observation in the signed chain. One POST is now all-or-nothing (the chain, step counters, decision counts and images are restored when any part is rejected).
  • Decisions keep what the agent chose. {action, target, reason, confidence} sent with a client-reported decision were dropped, so the record could not say what the agent decided or why. They are now fields of the decide entry (other unknown fields go in extra), and are hashed as the decision's text when no raw model output is sent. The tamper numbers above are unchanged: the fixes touch what is written, not how entries are chained or verified.