{"schema_version":"1","site":"https://decosa.ai","id":"evidence-runner","num":"140","name":"Capture audit evidence from your admin screens","tool_name":"Capture audit evidence from your admin screens","short":"Evidence runner","blurb":"For the compliance lead who screenshots 30 to 150 admin settings screens every quarter for SOC 2 or CMMC. Name the screens once (2-step verification, password policy, audit logging, your own product's admin pages). The agent finds each one, read-only, and code reads the settings you named; you approve those values as the baseline. Every quarter after that it runs with no model: it opens each screen from its saved route, re-reads every setting, captures it on a fresh load stamped with the URL, the time and who is signed in, and lists what changed since the baseline, all in a signed certificate. The browser session blocks every write, so it cannot change a setting. It runs on your side, in your own signed-in browser or inside your network.","status":"preview","labels":{"industry":["compliance-trust","software"],"job":["act","attest"],"input":["screen"],"deploy":["selfhost"],"status":"preview","output":["record","data"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["compliance-trust","software"],"runs_in":["selfhost"],"part_of":[],"built_from":["flight-recorder","signed-record"],"models":"Qwen3.8-27B finds each screen once (setup) and repairs routes when a console is redesigned; reading the settings, the quarterly replay, the stamps and the certificate are plain code","where":"On your side: a local helper attached to your own signed-in browser, or self-hosted in your network. The hosted demo runs on made-up consoles only","hardware":"The quarterly run needs only a CPU (headless Chromium). Setup and repairs call Qwen3.8-27B: 1x RTX PRO 6000 (96 GB) self-hosted, or the hosted gateway","final_artifact":"An evidence pack: a stamped PNG per screen, a manifest (screen, controls, URL, UTC time, SHA-256, matches or changed) and a signed decosa.test-run.v1 certificate.","self_host_first":true,"verification":{"receipt_coverage":"partial","summary":"A decosa.test-run.v1 certificate per run (every setting re-read by code, stamped screenshots hashed in, signed); receipt per model call at setup","manual_qa":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":7500,"p95_ms":10200,"runs":3,"receipts_per_run":0,"cost_per_run_usd":0},"selfhost":{"date":"2026-09-29","result":"pass","method":"fresh clone, venv, tests, then the CLI attached over CDP to a separate Chromium profile standing in for the admin's own browser, against a made-up console served over plain HTTP","notes":"Setup 3 of 3 screens right in 46 s on the local model server; next quarter 2 changes found in 2 s (exit 1), 0 writes reached the console, the stand-in browser's own tab left untouched."},"known_limits":["Made-up consoles only so far; no real tenant has been run.","Only the settings you name are asserted in the certificate; other settings on the same screen are compared and reported, not asserted.","The URL and time band is drawn onto each PNG by the runner (the certificate binds the file to its capture); it is not the browser's own address bar or the system clock.","Microsoft 365 and Entra admin centers: witnessed capture only (no agent navigation); use Microsoft Graph for settings.","Setup needs a person to approve the baseline values and to handle screens left for a person (pages with text aimed at AI agents, screens the agent could not find).","Hosted runs are demos on made-up consoles; real consoles run on your side."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Quarterly verdicts right on held-out consoles (no model)","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Four held-out made-up consoles; 9 / 9 changed screens caught, 0 false alarms."},{"name":"Quarterly verdicts right on development consoles","value":"19 / 19","unit":null,"n":19,"split":"dev","note":"7 / 7 changed screens caught."},{"name":"Setup: screen found and every named setting read right, held-out after fixes","value":"25 / 26","unit":null,"n":26,"split":"test","note":"Three held-out consoles on the hosted gateway, after four bugs found on the first look were fixed (a second look)."},{"name":"Setup, first look at two held-out consoles","value":"10 / 17","unit":null,"n":17,"split":"test","note":"Before the fixes; the failures exposed four bugs."},{"name":"Setup, first look at a console added after every fix","value":"8 / 10","unit":null,"n":10,"split":"test","note":"Both misses came from one engine bug (a menu link read as an action), fixed after this run."},{"name":"Setup, development consoles","value":"19 / 20","unit":null,"n":20,"split":"dev","note":null},{"name":"Wrong values read at setup","value":"0","unit":null,"n":null,"split":"test","note":"In every run; values code could not place were left for the person."},{"name":"Redesigned console handled (routes repaired)","value":"19 / 19","unit":null,"n":19,"split":"dev","note":null},{"name":"Writes that reached a console","value":"0","unit":null,"n":null,"split":"test","note":"Every run, including pages that asked agents to reset and save."}],"dataset":"Six made-up admin consoles written for this eval (identity admin, cloud console, an own-product admin with iframe pages, an endpoint manager with dialogs, code-hosting organisation settings, an HR and payroll system), two quarters of settings each, a redesigned variant of the two development consoles, planted text aimed at AI agents and session expiries.","held_out":false,"caveats":["The same author wrote the consoles and the runner; the consoles copy real shapes but are not real products.","Held-out consoles were looked at more than once: the first look exposed bugs that were fixed, so later numbers on them are second looks. The first-look numbers are listed separately.","No real tenant, no human reviewer timing; review time is not measured.","Small n per console (7-10 screens)."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/evidence-runner"},"stack":{"summary":"For the compliance lead or consultant who screenshots admin settings screens every quarter for SOC 2 or CMMC. Name the screens once; an agent finds each one, read-only, and code reads the settings you named; you approve those values as the baseline. Every quarter after that it runs with no model: saved routes reopen each screen, code re-reads every setting, each screen is captured on a fresh load and stamped, and changes since the baseline come first (with their direction, and any other setting on the same screen that moved). It runs on your side, in your own signed-in browser or inside your network, and its browser session blocks every write.","tagline":"Named admin settings screens captured every quarter with URL, time and user stamped, each setting re-read against the baseline you approved, in a signed certificate. Read-only.","deployment":"self-host-first","regulatory_note":"Read 28 Sep 2026 (public terms only; not legal advice): Google Workspace (Terms, last modified 2 Sep 2026; Acceptable Use Policy, 13 Oct 2025), the AWS Customer Agreement (14 Aug 2026) and Acceptable Use Policy (1 Jul 2021), and the Okta Master Subscription Agreement (Q1 FY26) say nothing against a customer automating read-only views of its own tenant, and each offers read-only APIs (Cloud Identity Policy API, the IAM credential report and a SecurityAudit role, okta.policies.read) that should come first. Microsoft's Product Terms bar using an Online Service \"to scrape or use other data extraction methods\", so for Microsoft 365 and Entra this tool offers only witnessed capture (a person clicks, it stamps and signs) and Microsoft Graph is the route for settings. Okta's MSA also bars using the service for others as a managed service: run the tool as the customer's own. Auditors ask for a URL and date stamp on screenshots (Secureframe, 16 Jul 2026) and for the completeness and accuracy of system-generated evidence (PCAOB AS 1105.10). This tool captures and compares; it does not say a control is effective, and it is not an assessment.","components":[{"id":"runner","role":"Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)","name":"decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12, Playwright with Chromium (BSD-3-Clause), Pillow","receipt_coverage":"partial","in_hosted_demo":null,"tiers":[],"alternative_to":null},{"id":"model","role":"Setup and repairs only: finds each named screen (read-only) and points at rows when code cannot place a setting","name":"Qwen3.8-27B","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27B","quant":"NVFP4","vram_gb":96,"memory_gb_estimate":null,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":null,"tiers":[],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · quarterly runs only, no GPU","summary":"Replays an approved baseline spec: saved routes, every setting re-read in code, stamped captures, certificate, evidence pack, offline verify. No model; setup must come from the standard tier or a colleague's approved spec.","components":["runner"],"hardware":"Any CPU","quality_evidence":[{"metric":"Quarterly verdicts right, six made-up consoles","value":"52 / 52 screens (33 / 33 on held-out consoles); 16 / 16 changed screens caught; 0 false alarms","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"}],"latency_note":"measured: about a second per screen","in_hosted_demo":true,"receipt_coverage":"none","receipt_note":"No model calls; the certificate is signed by the instance key.","hosting":null},{"id":"standard","label":"Standard · setup and repairs with Qwen3.8-27B","summary":"The agent finds each named screen read-only (two runs must agree) and repairs saved routes after a redesign; code reads the values; a person approves the baseline.","components":["runner","model"],"hardware":"1x RTX PRO 6000 96 GB (measured) or the hosted gateway","quality_evidence":[{"metric":"Setup, held-out consoles after fixes","value":"25 / 26 screens found with every named setting read right; 0 wrong values","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026 (hosted gateway run)"},{"metric":"Setup, first look at a console added after every fix","value":"8 / 10 (both misses from one bug, fixed after)","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"},{"metric":"Redesigned console, routes repaired","value":"19 / 19 screens","source":"decosa-api docs/evals/evidence-runner.md, 29 Sep 2026"}],"latency_note":"measured: under a minute per screen under shared load; a fraction of a cent per screen at list price","in_hosted_demo":true,"receipt_coverage":null,"receipt_note":"Setup decisions on the hosted gateway carry signed receipts.","hosting":null}],"alternates":[{"id":"api-export","label":"Official read-only API exports (Graph, Policy API, AWS, Okta)","components":["runner"],"hardware":"Any CPU","use":"The first route where an API covers a setting; not built into this tool yet (compliance platforms already do it).","status":"not built"}],"services":[{"name":"decosa-api (local helper or in-network service)","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon; build from source)","purpose":"The runner, the CLI (setup, approve, certify, verify, witness) and /testruns/verify."},{"name":"Qwen3.8-27B on vLLM (optional)","port":8114,"image":"vllm/vllm-openai:v0.29.0","purpose":"Setup and repairs, self-hosted; or use the hosted gateway with a key."}],"tools":[{"name":"Test-run certificate verifier","url":null,"license":"AGPL-3.0-or-later","purpose":"The certificate format (decosa.test-run.v1) and its verifier; the CMMC evidence map reads it as is."},{"name":"Google Cloud Identity Policy API (read-only settings)","url":"https://docs.cloud.google.com/identity/docs/concepts/supported-policy-api-settings","license":"Google terms","purpose":"The official route for Workspace security settings where it covers them."},{"name":"Microsoft Graph conditional access and authentication methods (read)","url":"https://learn.microsoft.com/en-us/graph/api/conditionalaccessroot-list-policies","license":"Microsoft terms","purpose":"The route for Microsoft 365 and Entra settings (no agent navigation there)."},{"name":"AWS IAM credential report","url":"https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_getting-report.html","license":"AWS terms","purpose":"System-generated evidence for IAM users and MFA."},{"name":"Okta Management API, Policies","url":"https://developer.okta.com/docs/api/openapi/okta-management/management/tag/Policy/","license":"Okta terms","purpose":"Read-only policy export with okta.policies.read."}],"hardware":[{"tier":"Any laptop or small VM (quarterly runs)","fits":true,"notes":"Headless Chromium on CPU; no GPU and no model calls."},{"tier":"1x RTX PRO 6000 96 GB (setup, self-hosted)","fits":true,"notes":"Qwen3.8-27B NVFP4; or the hosted gateway instead."}],"latency":[{"lane":"quarterly run, 9-10 screens, no model","typical_ms":7500,"source":"measured on our server 2026-09-29: 6.3-10.2 s per run over 6 made-up consoles (0.7-1.5 s per screen)"},{"lane":"setup, per screen (two agent runs must agree)","typical_ms":25000,"source":"measured on our server 2026-09-29 through the hosted gateway under shared load: per-screen p50 17-43 s across four consoles"}],"benchmark":null,"notes":["Measured only on made-up consoles that copy the shapes of real admin screens (identity admin, cloud console, a company's own product admin in iframes, an endpoint manager with dialogs, code-hosting organisation settings, an HR and payroll system). No real tenant yet.","Setup found the screen and read every named setting right on 19 of 20 development screens; on the first look at two held-out consoles 10 of 17 (four bugs found and fixed); after the fixes 25 of 26 on three held-out consoles; 8 of 10 on a sixth console run once at the end (both misses from one bug, since fixed). No wrong value was read in any run: what code cannot place is left for the person.","Quarterly runs gave the right verdict on every screen of all six consoles (52 of 52; 16 of 16 changed screens caught, no false alarms), with no model calls. A redesigned console: 19 of 19 screens handled (broken routes repaired by the agent, new values right, sent for approval).","No run changed a setting: in every run no write reached the console. A page carrying text aimed at AI agents stops the agent before the model reads it; that screen is captured by code only or left for the person.","Engine changes held MiniWoB held-out (181 of 252, main 180) and Gitea held-out (43 of 48, main 42)."]},"buyer_facts":[{"label":"Data retention","value":"Hosted demo: runs on made-up consoles are kept one hour for the session that made them, then deleted. Your own runs stay on your machine: the spec, the evidence packs and the signing key live where you run the helper."},{"label":"What leaves the box","value":"Quarterly runs: nothing (no model calls). Setup and repairs: screenshots and the element table of the admin screens the agent opens go to the model you choose (your own Qwen3.8-27B, or the hosted gateway with a key). Passwords never do: you sign in yourself."},{"label":"Read-only by construction","value":"The runner's tab blocks every write at the network level and no approval can release one; it refuses toggles, pickers and save-like clicks; every capture is a fresh load of the page."},{"label":"Where it runs","value":"On your side: a local helper attached to the browser you are already signed in to (it opens and gates one tab of its own), or inside your network. Your SSO, conditional access and device checks stay yours."},{"label":"Typical quarter","value":"A quarterly run takes seconds with no model calls; setup takes under a minute and a fraction of a cent per screen, once per console."},{"label":"What it will not do","value":"Say a control is effective, certify, or change a setting. It captures, compares with the baseline you approved, and signs."}],"data_handling":{"page":"/data#evidence-runner","self_host":{"level":"confidential","leaves":"content","summary":"Runs on your machine, but by default sends content to the services listed in external_calls."},"hosted":{"level":"operator-processed","demo_only":true,"summary":"Hosted demo on sample or public data only; self-host for real data.","gpus":"operator-contracted","third_parties":["The model you choose for setup and repairs"],"retention":"Hosted demo: runs on made-up consoles are kept one hour for the session that made them, then deleted. Your own runs stay on your machine: the spec, the evidence packs and the signing key live where you run the helper.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"The model you choose for setup and repairs (your own server, or the hosted gateway)","route":"both","sends":"content","what":"Screenshots and the element table of the admin screens the agent opens during setup or a repair; never passwords; never during quarterly runs.","default":"on","off":"Use a self-hosted model server (DECOSA_LLM_ROUTE=direct), or reuse an approved spec and run quarterly only."}]},"console":{"href":"/tools/finance/evidence-runner","input":"runner","lanes":[{"id":"setup","title":"Set up once: the agent finds each screen, you approve the baseline","kind":"list"},{"id":"quarter","title":"Every quarter: saved routes, every setting re-read, no model","kind":"list"},{"id":"capture","title":"Stamped captures (URL, UTC time, signed-in user, SHA-256)","kind":"list"},{"id":"record","title":"Evidence pack and signed decosa.test-run.v1 certificate","kind":"json"}],"samples":[{"n":1,"id":"keystone-quarterly","title":"Keystone quarterly","deep_link":"/tools/finance/evidence-runner?sample=1&autorun=0"},{"n":2,"id":"keystone-setup","title":"Keystone setup","deep_link":"/tools/finance/evidence-runner?sample=2&autorun=0"},{"n":3,"id":"northwind-quarterly","title":"Northwind quarterly","deep_link":"/tools/finance/evidence-runner?sample=3&autorun=0"},{"n":4,"id":"harbor-quarterly","title":"Harbor quarterly","deep_link":"/tools/finance/evidence-runner?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/evidence-runner-hosted.md","selfhost":"/prompts/evidence-runner-selfhost.md","assemble":"/prompts/evidence-runner-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/evidence-runner.zip","bundle_url":"https://decosa.ai/samples/evidence-runner.zip","folder":"/samples/evidence-runner/","expected":"/samples/evidence-runner/expected.json","files":["/samples/evidence-runner/expected.json"],"bytes":1245,"checks":["the hosted service offers the made-up consoles only","nine screens are captured","exactly the four planted screens changed","the minimum password length change is found (12 to 14)","the setting that flipped next to a named one is reported","every saved route reached its screen","the quarterly run makes no model calls","no write reached the console","the certificate verifies","a certificate with one check flipped no longer verifies"],"licence":"Synthetic: the console, its company and its settings were written for Decosa (decosa-api, AGPL-3.0-or-later). No real product, tenant or person.","about":"Keystone Admin is a made-up identity admin console served inside the runner's own browser (nothing leaves it). The approved baseline was made by a setup run on last quarter's settings; this quarter four named settings changed and one setting next to a named one flipped. The run opens every screen from its saved route with no model calls, re-reads each setting, captures every screen, and signs a certificate. It must report exactly the planted changes, reach every screen, let no write through, and the certificate must verify while a copy with one check flipped must not.","run":{"containers":"docker compose exec api python scripts/rehearse.py evidence-runner","checkout":"python scripts/rehearse.py evidence-runner --bundle evidence-runner.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py evidence-runner"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=evidence-runner","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"alternate-api-export","gpu_gb":0,"basis":null,"unknown":[]}],"mac":null},"links":{"page":"/tools/finance/evidence-runner","json":"/use-cases/evidence-runner.json","metrics":"/metrics/evidence-runner","console":"/tools/finance/evidence-runner","console_sample":"/tools/finance/evidence-runner?sample=1&autorun=0","stack":"/tools/finance/evidence-runner#stack","try_live":"/tools/finance/evidence-runner","watch":"/tools/finance/evidence-runner","build":"/tools/finance/evidence-runner#build","self_host":"/tools/finance/evidence-runner#self-host","prompts":{"hosted":"/prompts/evidence-runner-hosted.md","selfhost":"/prompts/evidence-runner-selfhost.md","assemble":"/prompts/evidence-runner-assemble.md","mac":null}}}