Skip to content
decosa

27 · Software and AI ops · Compliance and trust · live

Verified end-to-end test runs

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Verdict agrees with the scripted Playwright test, first run of each case52/52test splitn = 52
  • Verdict agrees, all runs120/120test splitn = 120
  • Seeded bugs and faulty users caught24/24test splitn = 24Each at the same step as the scripted test.
  • False fails on working builds (including the redesign)0/96test splitn = 96
  • Flaky cases (verdict changed between repeats)0/52test splitn = 52Every case ran 2-4 times; bounds the flip rate only loosely (roughly under 3% at 95% confidence).
  • Tampered certificates caught1,320/1,320test splitn = 1,320

Dataset

52 cases: 5 specs x 8 builds of a fixture shop app written for this eval (a good build, a redesign and 6 seeded bugs), plus 3 saucedemo.com specs x 4 public test users with known faults. Ground truth from hand-written Playwright scripts; 120 runs; 11 tamper alterations on every certificate.

Caveats

  • Small, and written by us: the fixture app, its bugs and the specs were written by the same author as the runner (the saucedemo faults are Sauce Labs').
  • Nothing was tuned on these cases (the action prompt is the flight recorder's, unchanged), but 52/52 shows it works on simple shop flows, not on a large product.
  • The model is language-only: canvas-heavy and cross-origin-iframe apps are not covered, and layout or visual regressions are out of scope.
  • Latency was measured while the shared GPU was saturated by other evals.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
86 s
Receipts
6
Model calls
n/a
Tokens
n/a
Cost per run
$0.002

Self-host verification

Verified on 25 Sep 2026: fresh clone, image built with WITH_BROWSER=1, compose up, sample against local model servers

Images build, the service starts, the sample passes on the good build (exit 0, 9 s) and fails on the seeded bug (exit 1, 19 s) against the already-running local Qwen3.8-27B vLLM (host network, no llm service started); a flipped assertion fails /testruns/verify; a private http target listed in DECOSA_TESTRUNS_TARGETS passes and an unlisted one is refused. Named volume for /data. Model-server startup itself not re-verified.

Rehearsal bundle: test-runs.zip (2 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Eval cases are simple shop flows written for it (plus Sauce Labs' public site); expect more stuck steps on complex apps.
  • The action model reads the element table, not pixels: canvas-heavy and cross-origin iframe apps are not supported.
  • Assertions read the DOM; layout and visual regressions are out of scope.
  • A run takes about 1.5 minutes when the shared GPU is busy (seconds when it is quiet); a scripted Playwright test is faster and free.
  • Hosted runs only reach the fixture app, saucedemo.com and domains you verified; hosted certificates are kept 24 hours.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI clientAGPL-3.0-or-later (the decosa_testrun CI client is Apache-2.0)
  • Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Headless browser that carries out the steps and reads the assertionsPlaywright 1.58 with Chromium headless shellApache-2.0 (Playwright); BSD-3-Clause (Chromium)

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · navigation-only specs, any CPU (1)
  • Navigation steps in the eval: 86 goto steps, all judged by the same assertion code as the agent stepsdecosa-api docs/evals/test-runs.md
Standard · agent steps, one 96 GB card (hosted demo) (4)
  • Verdict agrees with a hand-written Playwright test (52 cases: 5 specs x 8 fixture builds, 3 saucedemo specs x 4 users): 52/52 first runs; 120/120 over all runsdecosa-api docs/evals/test-runs.md, 2026-09-25
  • Seeded bugs and faulty users caught: 24/24 runs, each at the same step as the scripted test; 0/96 false fails on working builds, including a UI redesigndecosa-api docs/evals/test-runs-results.json
  • Flaky cases over 2-4 repeats: 0/52decosa-api docs/evals/test-runs.md
  • Tampered certificates caught (11 alterations): 1,320/1,320 with the issuer key pinneddecosa-api docs/evals/test-runs-results.json

How we measure · All tools