Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: verified end-to-end test runs (vertical 27)

Run on 25 Sep 2026 on our server: decosa-api branch the pre-release branch on 127.0.0.1:8482, gateway route (Qwen3.8-27B, every action receipted), signing with the live instance key. Script: scripts/testruns_eval.py; every run is listed in docs/evals/test-runs-results.json.

Question

Does a run driven by the agent (the model picks every click and keystroke from plain-language steps) reach the same verdict as a hand-written Playwright test of the same flow, on working builds and on builds with seeded bugs, and does the verdict stay the same when the run is repeated?

Set-up

  • 52 cases. 5 fixture specs x 8 builds of the Kiln & Co fixture app (good; redesign, which changes labels and product order but not behaviour; and 6 seeded bugs: subtotal ignores quantity, sign-up throws, promo takes 20% not 10%, cart badge never updates, Remove deletes the wrong line, Review order disabled), plus 3 saucedemo.com specs x 4 of its public test users (standard_user, and problem_user, locked_out_user, error_user, which Sauce Labs ships with known faults).
  • Ground truth. A hand-written Playwright script per spec (fixed data-test/id selectors, no model) performs the same steps, and the same assertion code (runner.evaluate + spec.judge) decides each step, stopping at the first failure. 12 of the 52 cases fail in the ground truth (6 fixture bugs, 6 saucedemo user faults); 40 pass. So the eval measures only whether the agent's actions reach the state a scripted test reaches; judging is identical code.
  • Runs. Every case once, again once more, and the 8 baseline cases (good build, standard_user) two more times: 120 runs, 4 at a time.
  • Tampering. 11 alterations applied to every certificate (1,320 trials), verified offline with the issuer key pinned.

Results

Result
Verdict agrees with the scripted test, first run of each case 52/52
Verdict agrees, all 120 runs 120/120
Seeded bugs and faulty users caught (runs whose truth is fail) 24/24, each at the same step as the scripted test
False fails on working builds (including the redesign) 0/96
Flaky cases (a verdict that changed between repeats of a case) 0/52 (every case ran 2-4 times)
Runs that ended in error 0
Certificates that verify 120/120
Tampered certificates caught 1,320/1,320 (flipped assertion, edited value read, assertion kind swapped, assertion deleted, step result flipped, verdict entry flipped, statement verdict flipped, build id changed, spec text edited, decision text edited, final screenshot changed)
Actions with a gateway-signed receipt 692/692
Model actions per run mean 5.8, max 12
Tokens and cost per run about 4,400 tokens (about 700 prompt tokens per action), $0.0015 at the gateway list price ($0.30 / $1.50 per million)
Certificate size median 72 KB with thumbnails, largest 154 KB
Time per action (model) median 12.8 s, 10-90% 9.9-21.9 s, with the shared model server saturated by other evals
Time per run median 86 s under that load; the same 3-step spec took 7.8 s on a quiet gateway (0.5-1.4 s per action) and 9-19 s self-hosted on the direct route

How steps ended: 218 when the assertions passed after an action, 86 plain navigations, 21 when the model said done, 1 at the action limit, 16 not run after an earlier failure. No guard fired (no spec asked for a final or payment button).

Honest limits

  • Small, and written by us. The fixture app, its bugs and all 8 specs were written for this eval, by the same author as the runner; the saucedemo faults are Sauce Labs'. The action prompt is the flight recorder's, unchanged, and nothing was tuned on these cases, but a 100% agreement on 52 cases says the approach works on simple shop flows, not that it will on a large product. Expect more max_actions and blocked steps on complex UIs.
  • Flake rate is 0 over 2-4 repeats per case. That bounds the per-run verdict flip rate only loosely (roughly under 3% with 95% confidence over the 120 runs). Model nondeterminism under load was not seen to change a verdict here.
  • The model is language-only. It reads the element table and page text, not pixels, so canvas-heavy and cross-origin-iframe apps need another observation source.
  • Assertions see the DOM. Layout and visual regressions are out of scope, and a step passes only on what it asserts.
  • Domain verification was tested with unit fakes, live DNS queries of public TXT records and live well-known fetches; no domain we own was put through the full claim-and-check flow.
  • Latency was measured while the shared GPU was saturated by other evaluation jobs; the per-run numbers above are dominated by queueing.