140 · Compliance and trust · Software and AI ops · preview
Capture audit evidence from your admin screens
Eval results
Not held outRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Quarterly verdicts right on held-out consoles (no model)33 / 33test splitn = 33Four held-out made-up consoles; 9 / 9 changed screens caught, 0 false alarms.
- Quarterly verdicts right on development consoles19 / 19dev (tuned on)n = 197 / 7 changed screens caught.
- Setup: screen found and every named setting read right, held-out after fixes25 / 26test splitn = 26Three held-out consoles on the hosted gateway, after four bugs found on the first look were fixed (a second look).
- Setup, first look at two held-out consoles10 / 17test splitn = 17Before the fixes; the failures exposed four bugs.
- Setup, first look at a console added after every fix8 / 10test splitn = 10Both misses came from one engine bug (a menu link read as an action), fixed after this run.
- Setup, development consoles19 / 20dev (tuned on)n = 20
- Wrong values read at setup0test splitIn every run; values code could not place were left for the person.
- Redesigned console handled (routes repaired)19 / 19dev (tuned on)n = 19
- Writes that reached a console0test splitEvery run, including pages that asked agents to reset and save.
Dataset
Six made-up admin consoles written for this eval (identity admin, cloud console, an own-product admin with iframe pages, an endpoint manager with dialogs, code-hosting organisation settings, an HR and payroll system), two quarters of settings each, a redesigned variant of the two development consoles, planted text aimed at AI agents and session expiries.
Caveats
- The same author wrote the consoles and the runner; the consoles copy real shapes but are not real products.
- Held-out consoles were looked at more than once: the first look exposed bugs that were fixed, so later numbers on them are second looks. The first-look numbers are listed separately.
- No real tenant, no human reviewer timing; review time is not measured.
- Small n per console (7-10 screens).
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 7.5 s
- Receipts
- 0
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0
Self-host verification
Verified on 29 Sep 2026: fresh clone, venv, tests, then the CLI attached over CDP to a separate Chromium profile standing in for the admin's own browser, against a made-up console served over plain HTTP
Setup 3 of 3 screens right in 46 s on the local model server; next quarter 2 changes found in 2 s (exit 1), 0 writes reached the console, the stand-in browser's own tab left untouched.
Rehearsal bundle: evidence-runner.zip (1 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Made-up consoles only so far; no real tenant has been run.
- Only the settings you name are asserted in the certificate; other settings on the same screen are compared and reported, not asserted.
- The URL and time band is drawn onto each PNG by the runner (the certificate binds the file to its capture); it is not the browser's own address bar or the system clock.
- Microsoft 365 and Entra admin centers: witnessed capture only (no agent navigation); use Microsoft Graph for settings.
- Setup needs a person to approve the baseline values and to handle screens left for a person (pages with text aimed at AI agents, screens the agent could not find).
- Hosted runs are demos on made-up consoles; real consoles run on your side.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)AGPL-3.0-or-later
- Setup and repairs only: finds each named screen (read-only) and points at rows when code cannot place a settingQwen3.8-27BApache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · quarterly runs only, no GPU (1)
- Quarterly verdicts right, six made-up consoles: 52 / 52 screens (33 / 33 on held-out consoles); 16 / 16 changed screens caught; 0 false alarmsdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026
Standard · setup and repairs with Qwen3.8-27B (3)
- Setup, held-out consoles after fixes: 25 / 26 screens found with every named setting read right; 0 wrong valuesdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026 (hosted gateway run)
- Setup, first look at a console added after every fix: 8 / 10 (both misses from one bug, fixed after)decosa-api docs/evals/evidence-runner.md, 29 Sep 2026
- Redesigned console, routes repaired: 19 / 19 screensdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026