27 · Software and AI ops · Compliance and trust · live
Verified end-to-end test runs
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Verdict agrees with the scripted Playwright test, first run of each case52/52test splitn = 52
- Verdict agrees, all runs120/120test splitn = 120
- Seeded bugs and faulty users caught24/24test splitn = 24Each at the same step as the scripted test.
- False fails on working builds (including the redesign)0/96test splitn = 96
- Flaky cases (verdict changed between repeats)0/52test splitn = 52Every case ran 2-4 times; bounds the flip rate only loosely (roughly under 3% at 95% confidence).
- Tampered certificates caught1,320/1,320test splitn = 1,320
Dataset
52 cases: 5 specs x 8 builds of a fixture shop app written for this eval (a good build, a redesign and 6 seeded bugs), plus 3 saucedemo.com specs x 4 public test users with known faults. Ground truth from hand-written Playwright scripts; 120 runs; 11 tamper alterations on every certificate.
Caveats
- Small, and written by us: the fixture app, its bugs and the specs were written by the same author as the runner (the saucedemo faults are Sauce Labs').
- Nothing was tuned on these cases (the action prompt is the flight recorder's, unchanged), but 52/52 shows it works on simple shop flows, not on a large product.
- The model is language-only: canvas-heavy and cross-origin-iframe apps are not covered, and layout or visual regressions are out of scope.
- Latency was measured while the shared GPU was saturated by other evals.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 86 s
- Receipts
- 6
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 25 Sep 2026: fresh clone, image built with WITH_BROWSER=1, compose up, sample against local model servers
Images build, the service starts, the sample passes on the good build (exit 0, 9 s) and fails on the seeded bug (exit 1, 19 s) against the already-running local Qwen3.8-27B vLLM (host network, no llm service started); a flipped assertion fails /testruns/verify; a private http target listed in DECOSA_TESTRUNS_TARGETS passes and an unlisted one is refused. Named volume for /data. Model-server startup itself not re-verified.
Rehearsal bundle: test-runs.zip (2 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Eval cases are simple shop flows written for it (plus Sauce Labs' public site); expect more stuck steps on complex apps.
- The action model reads the element table, not pixels: canvas-heavy and cross-origin iframe apps are not supported.
- Assertions read the DOM; layout and visual regressions are out of scope.
- A run takes about 1.5 minutes when the shared GPU is busy (seconds when it is quiet); a scripted Playwright test is faster and free.
- Hosted runs only reach the fixture app, saucedemo.com and domains you verified; hosted certificates are kept 24 hours.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI clientAGPL-3.0-or-later (the decosa_testrun CI client is Apache-2.0)
- Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Headless browser that carries out the steps and reads the assertionsPlaywright 1.58 with Chromium headless shellApache-2.0 (Playwright); BSD-3-Clause (Chromium)
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · navigation-only specs, any CPU (1)
- Navigation steps in the eval: 86 goto steps, all judged by the same assertion code as the agent stepsdecosa-api docs/evals/test-runs.md
Standard · agent steps, one 96 GB card (hosted demo) (4)
- Verdict agrees with a hand-written Playwright test (52 cases: 5 specs x 8 fixture builds, 3 saucedemo specs x 4 users): 52/52 first runs; 120/120 over all runsdecosa-api docs/evals/test-runs.md, 2026-09-25
- Seeded bugs and faulty users caught: 24/24 runs, each at the same step as the scripted test; 0/96 false fails on working builds, including a UI redesigndecosa-api docs/evals/test-runs-results.json
- Flaky cases over 2-4 repeats: 0/52decosa-api docs/evals/test-runs.md
- Tampered certificates caught (11 alterations): 1,320/1,320 with the issuer key pinneddecosa-api docs/evals/test-runs-results.json