{"schema_version":"1","site":"https://decosa.ai","id":"flight-recorder","num":"26","name":"Agent flight recorder","tool_name":"Prove what your agent did","short":"Flight recorder","blurb":"A signed, step-by-step record of what a browser or computer-use agent saw, decided and did. Replay it step by step; change one screenshot or one decision and verification names the step.","status":"live","labels":{"industry":["software","compliance-trust"],"job":["act","attest"],"input":["screen"],"deploy":["hosted","selfhost"],"status":"live","output":["record"],"data":["pii","confidential"],"hardware":"gpu-96","licence":"permissive","mode":"check"},"industries":["software","compliance-trust"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["flight-recorder","signed-record"],"models":"Qwen3.8-27B for decisions; any agent through the SDK","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB) for the decision model; the recorder and browser run on CPU","final_artifact":"A signed run record with thumbnails, receipts and guard events that anyone can re-verify.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipted decisions; signed, hash-chained run record","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":5293,"p95_ms":null,"runs":null,"receipts_per_run":8,"cost_per_run_usd":0.0025},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api (ba02fab), api image built from docker/api/Dockerfile (705 MB, no browser), compose from the assemble prompt with the llm service dropped and DECOSA_LLM_URL pointed at an already-running Qwen3.8-27B vLLM on the same box.","notes":"Verified on 2026-09-25: the image builds, the service starts, and the smoke test passes end to end against a local model server equivalent to the documented one; model-server startup itself was not re-verified. Key minting, the prompt's smoke script (verify, then fail at step 1 after a change), /decide (click on element 1, receipt status attested, 238 ms), 20-step timing (18.5 ms per step with a 480 KB screenshot, 3.9 ms hash-only) and the site's record viewer pointed at the box all worked. Worked around locally (fixed in the shared self-host pass): the compose health check calls curl, which the image does not have."},"known_limits":["Hosted demo sessions are limited per network each hour (the current number is in GET /healthz); the console also takes an API key.","A key has a per-minute request rate and a cap on open runs (GET /flight/info lists the limits). Close an abandoned run with DELETE /flight/runs/<id>, or see your runs with GET /flight/runs; the Python SDK waits out a 429.","Hosted decision latency depends on the shared model server: under a second per decision when quiet, several seconds when busy.","A restart of the hosted service stops a demo run in progress (shown as 'the demo run stopped').","The record proves what was reported and that it was not changed after signing; it does not prove a website did what it showed."],"nightly_covers":"one model decision, a sealed record and a tamper check through the API; no browser run"},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Genuine records that verify","value":"19/19","unit":null,"n":19,"split":"synthetic","note":"3 hosted-demo, 8 held-out SDK runs, 7 imported Jev runs, 1 synthetic OpenAI-shaped loop."},{"name":"Tampered copies caught, issuer key pinned","value":"339/339","unit":null,"n":339,"split":"synthetic","note":"22 kinds of alteration."},{"name":"Tampered copies caught without key pinning","value":"312/339","unit":null,"n":339,"split":"synthetic","note":"The 27 misses are chains rebuilt and re-signed with another key, which only pinning catches (by design)."},{"name":"Held-out agent runs reaching the expected outcome","value":"8/8","unit":null,"n":8,"split":"heldout","note":"4 tasks x 2 repeats; not a benchmark."},{"name":"MiniWoB++ success, production agent with guards","value":"25.9%","unit":null,"n":625,"split":"heldout","note":"95% CI 22.6-29.5%; 35.2% without the guards; no tuning for the bench."},{"name":"Mind2Web element accuracy / step success","value":"44.7% / 40.7%","unit":null,"n":300,"split":"heldout","note":"Raw model answers; 35% step success under the production value guard."},{"name":"Recording overhead per step with a screenshot (median)","value":"18.3","unit":"ms","n":40,"split":"synthetic","note":"8.4 ms hash-only."}],"dataset":"19 sealed flight records (hosted demo, held-out SDK runs on saucedemo.com and a fictional shop, imported Jev macOS runs) with 339 tamper trials; plus the demo agent on MiniWoB++ (125 tasks x 5 seeds) and a 300-step Mind2Web sample.","held_out":true,"caveats":["The record proves what was reported, not what happened: a client-reported step is only as honest as the agent.","Without key pinning, a record rebuilt and re-signed with another key verifies (0 of 27 caught).","Agent success of 8/8 on four held-out tasks is small; the decision prompt was adjusted on the three hosted demo tasks.","On public benchmarks the agent is demo-grade: 25.9% on MiniWoB++, failing on canvas, drag, custom widgets and unquoted values.","MiniWoB++ is tiny and synthetic and Mind2Web is offline; no live multi-page benchmark (WebArena) was run.","Decision latency was measured under heavy shared load."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/flight-recorder"},"stack":{"summary":"Your agent posts each step to the recorder: the page or screen it saw (a screenshot hash and thumbnail, and the element table or accessibility tree the model was shown), the model's decision, the action and the result. Every entry is appended to a hash chain; when the run ends it is sealed with an Ed25519 signature. Decisions can run on Qwen3.8-27B through our gateway, and then the signed receipt for that exact output sits inside the step. Guards run in code before anything is executed: a typed value must come from the task, and an order, payment or deletion waits for a person. A viewer replays the run step by step, and verification in the browser names the step where a screenshot, a decision or a guard was changed. It is for teams running agents on back-office work who need evidence after an incident, in a customer dispute or for an audit.","tagline":"A signed, step-by-step record of what a browser or computer-use agent saw, decided and did, that anyone can re-check.","deployment":"hosted-or-self-host","regulatory_note":"Checked 2026-09-25. EU AI Act: Article 12 requires high-risk AI systems to be able to record events automatically (logs) over their lifetime, and Articles 19 and 26(6) require providers and deployers to keep those logs for at least six months unless other law says otherwise. That duty applies only to high-risk systems (Annex I products, and Annex III uses such as hiring, credit scoring or access to essential services); after the AI Omnibus (Council approval 29 June 2026) Annex III duties apply from 2 December 2027 and Annex I from 2 August 2028. Most browser agents doing back-office work are not high-risk systems, and for them this record is useful evidence, not a legal requirement. Article 12 does not require signatures, hash chains or screenshots, and the harmonised logging standards (prEN 18229-1, ISO/IEC DIS 24970) are still drafts, so no product can claim conformity with them yet; this one does not. Screenshots of back-office screens usually contain personal data: under the GDPR, minimisation and storage limits apply, so use hash-only mode (images stay with you) or self-host. The hosted demo keeps runs for 24 hours and only visits a fictional shop. In a dispute a signature shows the record was not changed after signing and which key signed it; it does not show that what an agent reported about itself was true, and its weight as evidence is for the court or arbitrator. Model licence: Apache-2.0 (Qwen3.8-27B). Not legal advice.","components":[{"id":"recorder","role":"Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)","name":"decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK","hf_repo":null,"license":"AGPL-3.0-or-later (the SDK, the flight record format and its verifier are Apache-2.0)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.11+, FastAPI; Pillow for thumbnails; SDK is standard library only","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"decider","role":"Decision model: picks the next action from a numbered element table (never coordinates, never free text)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding (3 draft tokens)","vram_gb":57.6,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (OpenAI-compatible), called through our gateway's metered route; served language-only, so decisions read the element table, not pixels","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"browser","role":"Headless browser for the hosted demo and the Playwright adapter","name":"Playwright 1.58 with Chromium headless shell","hf_repo":null,"license":"Apache-2.0 (Playwright); BSD-3-Clause (Chromium)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"playwright==1.58.0, chromium_headless_shell-1208","receipt_coverage":"none","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"jev","role":"Typed decisions on Apple silicon (Jev-compatible local decision server), for the macos-harness adapter","name":"DiffusionGemma 26B-A4B (MLX, OptiQ 4-bit)","hf_repo":"mlx-community/diffusiongemma-26B-A4B-it-OptiQ-4bit","license":"Apache-2.0","params":"26B","quant":"OptiQ 4-bit (MLX)","vram_gb":18.5,"memory_gb_estimate":null,"engine":"MLX 0.32.2 + mlx-optiq 0.5.12, POST /v1/systemone (Saik0s/diffusiongemma-jev-macos at 5d89d73)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · record only, any CPU","summary":"Your agent and your model; the recorder chains, seals and verifies. Decisions from another model are recorded with their hash but carry no receipt.","components":["recorder"],"hardware":"Any Linux or macOS machine with Python 3.11+","quality_evidence":[{"metric":"Genuine records that verify","value":"19/19 (3 hosted demo runs, 8 SDK agent runs, 7 imported Jev harness runs, 1 computer-use loop)","source":"decosa-api docs/evals/flight-recorder.md, 2026-09-25"},{"metric":"Tampered copies caught (22 kinds of alteration)","value":"339/339 with the issuer key pinned; 312/339 without (the 27 were rebuilt and re-signed with another key, which only pinning can catch)","source":"decosa-api docs/evals/flight-recorder-results.json"},{"metric":"Overhead per step","value":"18 ms and 15 KB of record with a thumbnail; 8 ms and 3.3 KB hash-only","source":"decosa-api docs/evals/flight-recorder.md (40-step benchmark, local HTTP)"}],"latency_note":"measured: milliseconds per step on the recorder; the rest is your agent","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"The record is attested by the box's own key; decisions from a model outside this server are not receipted.","hosting":null},{"id":"standard","label":"Standard · receipted decisions, one 96 GB card (hosted demo)","summary":"Qwen3.8-27B makes each decision through /decide, with a gateway receipt inside the step and guards before anything runs.","components":["recorder","decider","browser"],"hardware":"1× RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Held-out agent tasks reaching the expected outcome (2 on saucedemo.com, 2 on the demo shop, 2 runs each; 2 expected a guard stop)","value":"8/8","source":"decosa-api docs/evals/flight-recorder.md; the prompt was adjusted on the three demo tasks, not these"},{"metric":"Decisions with a gateway-signed receipt","value":"83/83","source":"decosa-api docs/evals/flight-recorder.md"},{"metric":"Guard stops on an order or payment button","value":"3/3 runs whose task asked to place or finish an order stopped before the click","source":"decosa-api docs/evals/flight-recorder.md"},{"metric":"MiniWoB++ success, the agent on its own (125 tasks x 5 seeds, BrowserGym task classes)","value":"25.9% (95% CI 22.6-29.5%) with the guards; 35.2% (31.6-39.0%) without them; 36.7% on the form-and-button tasks, 4% on canvas, drag and slider tasks","source":"decosa-api docs/evals/computer-use-bench.md, 2026-09-25; production prompt, no tuning"},{"metric":"Mind2Web, next-step accuracy on real websites (300 test steps, top-50 candidates)","value":"element 44.7% (39.1-50.3%), step success 40.7% (35.3-46.3%) before the guards; 35% after them","source":"decosa-api docs/evals/computer-use-bench.md, 2026-09-25"}],"latency_note":"measured: seconds per decision while the shared GPU was saturated; about a minute for a short run","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every decision made through /decide is gateway-receipted and bound to its output hash in the chain.","hosting":null}],"alternates":[{"id":"jev","label":"Typed decisions on Apple silicon (Jev-compatible)","components":["jev"],"hardware":"Apple silicon with 32 GB or more","use":"DiffusionGemma as a local decision server for the macOS harness: fast single choices, weak at ordered sequences. Not a hosted model, so no receipts.","status":"self-host only"}],"services":[{"name":"decosa-api (flight routes)","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"POST /flight/runs, /flight/runs/{id}/steps, /decide, /seal; POST /flight/verify; POST /flight/demo (SSE); GET /flight/info, /flight/tasks, /flight/sdk/<file>."},{"name":"Decision model (vLLM, or our gateway)","port":8114,"image":"vllm/vllm-openai:v0.29.0","purpose":"Qwen3.8-27B for /decide. Not needed when your agent brings its own model."}],"tools":[{"name":"decosa_flight.py (SDK) and adapters: playwright_agent.py, jev_macos.py, openai_cu.py","url":null,"license":"Apache-2.0","purpose":"Served at GET /flight/sdk/<file>. Record any agent over HTTP with a dk_ key; hash-only mode sends no images. The Jev adapter imports a macos-harness run directory or hooks its loop; the OpenAI adapter wraps computer_call and computer_call_output."},{"name":"Harbor Supply (fictional demo shop)","url":null,"license":"Apache-2.0","purpose":"Six static pages served to the demo browser by request interception. Nothing is sold; the completion checks read its session state."},{"name":"Swag Labs demo site (saucedemo.com)","url":"https://www.saucedemo.com","license":"Public test site by Sauce Labs","purpose":"Held-out eval tasks only (public test credentials printed on its login page). The hosted demo never visits it."},{"name":"scripts/flight_eval.py","url":null,"license":"Apache-2.0","purpose":"Held-out agent runs, overhead benchmark and the tamper set (decosa-api)."}],"hardware":[{"tier":"Any CPU, no GPU","fits":true,"notes":"Recording, sealing and verification: 18 ms per step with a screenshot, 8 ms hash-only (local HTTP, measured)."},{"tier":"1× RTX PRO 6000 96 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 for receipted decisions; measured on our server."},{"tier":"Apple silicon, 32 GB or more","fits":true,"notes":"DiffusionGemma MLX 4-bit for the Jev macos-harness: about 18.5 GB peak per request (measured on an M3 Ultra). Its decisions are not receipted."}],"latency":[{"lane":"record one step (observation with a 1280x800 screenshot, decision, action, result), local HTTP","typical_ms":18,"source":"measured on our server 2026-09-25: median of 40, p95 20 ms; 8 ms median in hash-only mode"},{"lane":"one receipted decision (Qwen3.8-27B through the gateway)","typical_ms":7811,"source":"measured on our server 2026-09-25: median of 83 decisions, 10-90% 6.7-12.8 s, while the shared model server had 20-30 requests running and 50-70 queued"},{"lane":"hosted demo run, 8 steps (cheapest tool to review)","typical_ms":80600,"source":"measured on our server 2026-09-25 (run fr-f2fbde7758536e72), same GPU load"},{"lane":"verify a 44-entry record with 11 thumbnails","typical_ms":5,"source":"measured on our server 2026-09-25 (server-side); the browser check runs the same algorithm with WebCrypto"}],"benchmark":{"title":"How well does the agent do on its own?","intro":"The recorder's numbers above measure the record. This measures the demo agent: Qwen3.8-27B choosing actions from the numbered element table, on two public benchmarks, with the production prompt and no tuning. 25 Sep 2026, 95% intervals in brackets.","rows":[{"label":"MiniWoB++, 125 small web tasks x 5 seeds, production agent with guards","value":"25.9% (22.6-29.5%)","detail":"36.7% on tasks built from links, buttons and inputs; 4% on canvas, shape and colour tasks; 4% on drag, slider and keyboard tasks"},{"label":"Same, without the guards","value":"35.2% (31.6-39.0%)","detail":"The 9-point gap is almost all the value guard: it refused to type dates, sums and other values the task did not quote"},{"label":"Same weights reading the screenshot instead of the table (not served today)","value":"55.8% (51.9-59.7%)","detail":"Pixels only, coordinate clicks, no guards; 0.75 s per decision on a private server"},{"label":"Mind2Web, 300 recorded steps on real websites: right element / right element and operation","value":"44.7% / 40.7%","detail":"Brackets 39.1-50.3% and 35.3-46.3%. Under the production value guard, 35% of steps would go through"},{"label":"Decision time, production agent","value":"0.59 s median","detail":"p90 7.1 s when the shared model server was busy"}],"points":[{"heading":"What the element table means","text":"The model never sees the page. It gets a numbered list of the links, buttons, inputs and selects it can act on, and picks one. That makes each choice checkable and recordable, and it is why it does well on ordinary forms. It is also a crutch: anything not in the list does not exist for the agent. In 20% of MiniWoB episodes the list was empty at the first step, because the clickable things were plain spans and divs."},{"heading":"Where it fails","text":"Canvas and drawn shapes, colours, drag and sliders, custom widgets such as date pickers, and content inside iframes, which the extractor does not enter. Long, exploratory sequences fail too: an eight-step flight booking scored 0 of 5 with every setup, and paging through tabs to find a link often loops until the step limit. In most of these cases it says it is blocked rather than guessing."},{"heading":"Why the guards matter","text":"Without them, the model said it was done in 82 of 625 runs (13%) when the task had not finished, and it typed values nobody gave it. With them, a run is only marked successful when the page agrees, and anything typed comes from the task or the caller's allowed values. That costs capability, and the table shows how much. Pass the values a task needs as allowed_values rather than turning the guard off."},{"heading":"Verdict","text":"Demo-grade on its own. It is usable for narrow, form-shaped jobs on sites with real links and inputs, when the values are supplied and a completion check is written for the task, as in the hosted demo and the 8/8 held-out runs. It is not a general web agent. The vision result shows where the gain is: a hybrid that uses the table when it has the target and the screenshot when it does not, with the same guards."}],"source":"decosa-api docs/evals/computer-use-bench.md and computer-use-bench.json (every episode and model answer), 2026-09-25. MiniWoB++ (MIT) through BrowserGym (Apache-2.0); Mind2Web (CC BY 4.0) with the Multimodal-Mind2Web test subset (OpenRAIL). WebArena was not run: it needs six self-hosted sites."},"notes":["The record proves what was reported to the recorder, in what order and when, which model made each receipted decision, and that nothing changed after signing. It does not prove that a website really did what it displayed, or that a step an agent reported about itself is true. Steps observed by the server's own browser are marked \"seen by our browser\".","The model's \"done\" never counts as success. A completion check reads the page itself; when it fails, a note goes on the record and the agent continues. The eval's agent runs reached the expected outcome in 8 of 8 held-out runs.","Someone who holds the signing key can rebuild and re-sign a whole record; someone who does not can only produce a record signed by a different key. Pin the issuer's key (GET /attest/signing-key) and verification catches that too. Keep the key off the machine the agent runs on if the agent's operator is the party you need to hold to account.","Hash-only mode keeps screenshots on your side: the record holds their sha256, so you can later prove a stored screenshot is the one the agent saw without ever sending it."]},"buyer_facts":[{"label":"Data retention","value":"Hosted: runs and screenshots are kept 24 hours (DECOSA_FLIGHT_TTL_S), then deleted. Self-host: you set the TTL; sealed records verify offline, so archive them yourself."},{"label":"What leaves the box","value":"Hosted: whatever the agent posts (screenshots, page text, decisions), unless the SDK runs with send_images=False, which sends only screenshot hashes. Self-host: nothing, unless you point DECOSA_LLM_ROUTE at a gateway."},{"label":"Inputs and limits","value":"Screenshots as JPEG, PNG or WebP up to 3 MB each (4 MB of images per run; the server keeps a 320 px thumbnail and the hash), page text up to 16,000 characters per step, up to 300 steps per run."},{"label":"Logs","value":"The service log carries method, path, status and timing only; task text and screenshots are not logged (checked on the self-host box)."}],"data_handling":{"page":"/data#flight-recorder","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Hosted: runs and screenshots are kept 24 hours (DECOSA_FLIGHT_TTL_S), then deleted. Self-host: you set the TTL; sealed records verify offline, so archive them yourself.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/developer/flight-recorder","input":"flight","lanes":[{"id":"steps","title":"Steps","kind":"list"},{"id":"guards","title":"Guard events","kind":"list"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"cheapest-tool","title":"Cheapest tool","deep_link":"/tools/developer/flight-recorder?sample=1&autorun=0"},{"n":2,"id":"swap-cart","title":"Swap cart","deep_link":"/tools/developer/flight-recorder?sample=2&autorun=0"},{"n":3,"id":"place-order","title":"Place order","deep_link":"/tools/developer/flight-recorder?sample=3&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/flight-recorder-hosted.md","selfhost":"/prompts/flight-recorder-selfhost.md","assemble":"/prompts/flight-recorder-assemble.md","mac":"/prompts/flight-recorder-mac.md"},"rehearsal":{"bundle":"/samples/flight-recorder.zip","bundle_url":"https://decosa.ai/samples/flight-recorder.zip","folder":"/samples/flight-recorder/","expected":"/samples/flight-recorder/expected.json","files":["/samples/flight-recorder/expected.json","/samples/flight-recorder/inputs/cart-elements.json","/samples/flight-recorder/inputs/cart-text.txt","/samples/flight-recorder/inputs/cart.jpg","/samples/flight-recorder/inputs/review-elements.json","/samples/flight-recorder/inputs/review-history.json","/samples/flight-recorder/inputs/review-text.txt","/samples/flight-recorder/inputs/review.jpg","/samples/flight-recorder/inputs/run.json","/samples/flight-recorder/inputs/seal.json","/samples/flight-recorder/inputs/step-1-action.json"],"bytes":39879,"checks":["on the cart page the model chooses to click Checkout","on the review page the Place order click is caught by the needs_approval guard","the agent is stopped and the run escalated to a person","the guard event is chained in the sealed record","the sealed record verifies","both screenshots are in the record and match their recorded hashes","a record whose recorded click target was changed no longer verifies","verification points at step 1, where the change was made","every model call has a signed receipt"],"licence":"Synthetic: Harbor Supply is a fictional demo shop shipped with decosa-api (AGPL-3.0-or-later); the screenshots were rendered from its pages. No real shop, orders or people.","about":"A browser agent is asked to buy a mug on Harbor Supply, a fictional demo shop. Two observations (screenshot, element table and page text) go to the receipted decision model: on the cart page it must choose an action, and on the order review page the Place order click must be escalated by the needs_approval guard instead of executed. The run is sealed; the record must verify, and a copy whose recorded click target was changed must fail at step 1.","run":{"containers":"docker compose exec api python scripts/rehearse.py flight-recorder","checkout":"python scripts/rehearse.py flight-recorder --bundle flight-recorder.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py flight-recorder"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=flight-recorder","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"alternate-jev","gpu_gb":18.5,"basis":"stack","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/flight-recorder","json":"/use-cases/flight-recorder.json","metrics":"/metrics/flight-recorder","console":"/tools/developer/flight-recorder","console_sample":"/tools/developer/flight-recorder?sample=1&autorun=0","stack":"/tools/developer/flight-recorder#stack","try_live":"/tools/developer/flight-recorder","watch":"/tools/developer/flight-recorder","build":"/tools/developer/flight-recorder#build","self_host":"/tools/developer/flight-recorder#self-host","prompts":{"hosted":"/prompts/flight-recorder-hosted.md","selfhost":"/prompts/flight-recorder-selfhost.md","assemble":"/prompts/flight-recorder-assemble.md","mac":"/prompts/flight-recorder-mac.md"}}}