{"schema_version":"1","site":"https://decosa.ai","id":"test-runs","num":"27","name":"Verified end-to-end test runs","tool_name":"Run an end-to-end browser test","short":"Test runs","blurb":"Write a test as plain-language steps with assertions. A browser agent carries out the steps; pass or fail comes only from the assertions, checked in code. Each run ends in a signed certificate you can keep as release evidence, and a failed assertion fails the build.","status":"live","labels":{"industry":["software","compliance-trust"],"job":["act","attest"],"input":["screen"],"deploy":["hosted","selfhost"],"status":"live","output":["record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive","mode":"check"},"industries":["software","compliance-trust"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["flight-recorder","signed-record"],"models":"Qwen3.8-27B for actions; assertions in code","where":"Hosted (verified domains) or self-host","hardware":"1× RTX PRO 6000 (96 GB) for the action model; the runner, browser and verifier run on CPU","final_artifact":"A signed run certificate: spec, target, build, every assertion's value and a verdict anyone can re-derive.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipted actions; signed run certificate whose verdict recomputes","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":86054,"p95_ms":null,"runs":null,"receipts_per_run":6,"cost_per_run_usd":0.0015},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, image built with WITH_BROWSER=1, compose up, sample against local model servers","notes":"Images build, the service starts, the sample passes on the good build (exit 0, 9 s) and fails on the seeded bug (exit 1, 19 s) against the already-running local Qwen3.8-27B vLLM (host network, no llm service started); a flipped assertion fails /testruns/verify; a private http target listed in DECOSA_TESTRUNS_TARGETS passes and an unlisted one is refused. Named volume for /data. Model-server startup itself not re-verified."},"known_limits":["Eval cases are simple shop flows written for it (plus Sauce Labs' public site); expect more stuck steps on complex apps.","The action model reads the element table, not pixels: canvas-heavy and cross-origin iframe apps are not supported.","Assertions read the DOM; layout and visual regressions are out of scope.","A run takes about 1.5 minutes when the shared GPU is busy (seconds when it is quiet); a scripted Playwright test is faster and free.","Hosted runs only reach the fixture app, saucedemo.com and domains you verified; hosted certificates are kept 24 hours."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Verdict agrees with the scripted Playwright test, first run of each case","value":"52/52","unit":null,"n":52,"split":"test","note":null},{"name":"Verdict agrees, all runs","value":"120/120","unit":null,"n":120,"split":"test","note":null},{"name":"Seeded bugs and faulty users caught","value":"24/24","unit":null,"n":24,"split":"test","note":"Each at the same step as the scripted test."},{"name":"False fails on working builds (including the redesign)","value":"0/96","unit":null,"n":96,"split":"test","note":null},{"name":"Flaky cases (verdict changed between repeats)","value":"0/52","unit":null,"n":52,"split":"test","note":"Every case ran 2-4 times; bounds the flip rate only loosely (roughly under 3% at 95% confidence)."},{"name":"Tampered certificates caught","value":"1,320/1,320","unit":null,"n":1320,"split":"test","note":null}],"dataset":"52 cases: 5 specs x 8 builds of a fixture shop app written for this eval (a good build, a redesign and 6 seeded bugs), plus 3 saucedemo.com specs x 4 public test users with known faults. Ground truth from hand-written Playwright scripts; 120 runs; 11 tamper alterations on every certificate.","held_out":true,"caveats":["Small, and written by us: the fixture app, its bugs and the specs were written by the same author as the runner (the saucedemo faults are Sauce Labs').","Nothing was tuned on these cases (the action prompt is the flight recorder's, unchanged), but 52/52 shows it works on simple shop flows, not on a large product.","The model is language-only: canvas-heavy and cross-origin-iframe apps are not covered, and layout or visual regressions are out of scope.","Latency was measured while the shared GPU was saturated by other evals."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/test-runs"},"stack":{"summary":"You write a test as YAML: plain-language steps ('Add one Linen Apron to the cart') and, for each step, assertions (text present, URL, element count, value equals). A headless browser carries out each step; Qwen3.8-27B picks the clicks and typing from the page's element table, with a signed receipt per action, and never sees the assertions. Pass or fail comes only from the assertions, read from the live page and judged in code, and the first failed step ends the run. Every action, every assertion's expected and actual value, and the verdict are chained and signed into a run certificate that names the spec, the target and the build; anyone can verify it and re-derive the verdict. A CLI runs specs in CI, fails the build on a failed assertion and keeps the certificate as change-management evidence. For small SaaS teams without QA staff, and teams that have to show an auditor that a release was tested.","tagline":"Plain-language browser tests judged only by assertions in code, each run sealed into a signed certificate you can keep as release evidence.","deployment":"hosted-or-self-host","regulatory_note":"Checked 2026-09-25. SOC 2 (AICPA Trust Services Criteria, CC8.1) and ISO/IEC 27001:2022 (Annex A 8.29 security testing in development and acceptance, 8.32 change management) expect evidence that changes were tested and approved before release; neither requires a particular tool, a signature or a hash chain. A run certificate is one piece of evidence an auditor can sample and re-check; it does not make a change process compliant, does not record who approved the release, and covers only what the spec asserts. Testing a website without the owner's permission can be unlawful (for example under the US Computer Fraud and Abuse Act or the UK Computer Misuse Act); the hosted service therefore only visits its own fixture app, Sauce Labs' public test site and domains whose owner proved control with a DNS TXT record or a /.well-known file. Use test accounts and test data: screenshots of a staging site can contain personal data, and under the GDPR minimisation and storage limits apply (hosted certificates are kept 24 hours; self-host keeps everything on your box). Passwords go in run secrets and never reach the model or the certificate. Model licence: Apache-2.0 (Qwen3.8-27B). Not legal advice.","components":[{"id":"runner","role":"Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)","name":"decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client","hf_repo":null,"license":"AGPL-3.0-or-later (the decosa_testrun CI client is Apache-2.0)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.11+, FastAPI; PyYAML for YAML specs; Pillow for thumbnails; CI client is standard library only","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"decider","role":"Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding (3 draft tokens)","vram_gb":57.6,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (OpenAI-compatible), called through our gateway's metered route; served language-only, so decisions read the element table, not pixels","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"browser","role":"Headless browser that carries out the steps and reads the assertions","name":"Playwright 1.58 with Chromium headless shell","hf_repo":null,"license":"Apache-2.0 (Playwright); BSD-3-Clause (Chromium)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"playwright==1.58.0, chromium_headless_shell-1208","receipt_coverage":"none","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"vision-actor","role":"Vision action model for canvas and iframe apps (controls not in the DOM)","name":"Holo-3.1-35B-A3B (NVFP4)","hf_repo":"Hcompany/Holo-3.1-35B-A3B-NVFP4","license":"Apache-2.0","params":"35B","quant":"NVFP4 (23.7 GB of weights)","vram_gb":null,"memory_gb_estimate":24,"engine":"vLLM","receipt_coverage":"none","in_hosted_demo":false,"tiers":["alternates"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · navigation-only specs, any CPU","summary":"Every step is a goto with assertions: no model, no GPU, still a signed certificate. For smoke checks of pages, not flows.","components":["runner","browser"],"hardware":"Any Linux machine with 2 GB RAM per browser","quality_evidence":[{"metric":"Navigation steps in the eval","value":"86 goto steps, all judged by the same assertion code as the agent steps","source":"decosa-api docs/evals/test-runs.md"}],"latency_note":"estimate: a goto step and its assertions take under a second plus page load; no model wait","in_hosted_demo":true,"receipt_coverage":"partial","receipt_note":"No model calls, so no receipts; the certificate is attested by the box's own key.","hosting":null},{"id":"standard","label":"Standard · agent steps, one 96 GB card (hosted demo)","summary":"Qwen3.8-27B carries out the plain-language steps through the flight recorder's /decide, a receipt per action; assertions judge.","components":["runner","decider","browser"],"hardware":"1× RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Verdict agrees with a hand-written Playwright test (52 cases: 5 specs x 8 fixture builds, 3 saucedemo specs x 4 users)","value":"52/52 first runs; 120/120 over all runs","source":"decosa-api docs/evals/test-runs.md, 2026-09-25"},{"metric":"Seeded bugs and faulty users caught","value":"24/24 runs, each at the same step as the scripted test; 0/96 false fails on working builds, including a UI redesign","source":"decosa-api docs/evals/test-runs-results.json"},{"metric":"Flaky cases over 2-4 repeats","value":"0/52","source":"decosa-api docs/evals/test-runs.md"},{"metric":"Tampered certificates caught (11 alterations)","value":"1,320/1,320 with the issuer key pinned","source":"decosa-api docs/evals/test-runs-results.json"}],"latency_note":"measured: seconds per action and a minute or two per run with the shared GPU saturated; seconds for a short run on a quiet gateway","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every action is gateway-receipted and bound to its output hash in the chain.","hosting":null}],"alternates":[{"id":"vision-actions","label":"Vision actions for canvas and iframe apps","components":["vision-actor"],"hardware":"1x 96 GB card (estimate)","use":"Try the standard decider's own vision tower first (Qwen3.8-27B, same weights; screenshots took our MiniWoB++ run from 25.9% to 55.8%), then a dedicated grounder such as Holo-3.1 when pixel grounding is needed. Needs image-input receipts on the gateway.","status":"not served"}],"services":[{"name":"decosa-api (test-run routes)","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"POST /testruns/runs (SSE or JSON), /testruns/verify, /testruns/spec/check, /testruns/domains and /domains/check; GET /testruns/info, /samples, /runs/{id}, /sdk/decosa_testrun.py. Build with WITH_BROWSER=1."},{"name":"Action model (vLLM, or our gateway)","port":8114,"image":"vllm/vllm-openai:v0.29.0","purpose":"Qwen3.8-27B for the actions. Navigation-only specs need no model."}],"tools":[{"name":"decosa_testrun.py (CI client)","url":null,"license":"Apache-2.0","purpose":"Served at GET /testruns/sdk/decosa_testrun.py. Runs a spec file, writes and verifies the certificate, exits 0 pass / 1 fail / 2 error; --repeat writes a signed flake report."},{"name":"Kiln & Co (fictional fixture app)","url":null,"license":"Apache-2.0","purpose":"Seven static pages served to the browser by request interception, in 8 builds: good, a redesign and 6 seeded bugs. Nothing is sold."},{"name":"Swag Labs demo site (saucedemo.com)","url":"https://www.saucedemo.com","license":"Public test site by Sauce Labs","purpose":"Built-in public target; its test users (standard_user, problem_user, locked_out_user, error_user) and password are printed on its login page."},{"name":"scripts/testruns_eval.py","url":null,"license":"Apache-2.0","purpose":"Agent verdicts against hand-written Playwright tests, repeats for flake rate, and the tamper set (decosa-api)."}],"hardware":[{"tier":"Any CPU, no GPU","fits":true,"notes":"The runner and browser: about 2 GB RAM per concurrent browser. Specs made only of goto steps and assertions need nothing else."},{"tier":"1× RTX PRO 6000 96 GB","fits":true,"notes":"Qwen3.8-27B NVFP4 for the actions; measured on our server."},{"tier":"1× RTX 5090 32 GB","fits":true,"notes":"Qwen3.8-27B with a shorter context (about 700 prompt tokens per action is well inside it). Not measured here."}],"latency":[{"lane":"one action (Qwen3.8-27B through the gateway)","typical_ms":12793,"source":"measured on our server 2026-09-25: median of 692 actions, 10-90% 9.9-21.9 s, with the shared model server saturated by other evals"},{"lane":"one action on a quiet gateway","typical_ms":900,"source":"measured on our server 2026-09-25: 0.5-1.4 s per action on a 3-step run"},{"lane":"whole run (3-4 test steps, about 6 actions)","typical_ms":86054,"source":"measured on our server 2026-09-25: median of 120 eval runs under that load; 7.8 s on a quiet gateway, 9-19 s self-hosted on the direct route"},{"lane":"verify a certificate","typical_ms":10,"source":"estimate from the flight recorder's 5 ms for a 44-entry record; certificates here have 17-60 entries"}],"benchmark":null,"notes":["The model chooses actions only. It is never shown the assertions, and a 'done' that the assertions contradict gets one generic 'not finished' note, never what is checked, so the agent cannot steer towards a pass.","A step ends when its assertions pass after an action (unless they already passed before the step), when the model says done, or after max_actions. The first failed step ends the run; later steps are marked not run.","Verdict error means the run could not be carried out (start page, model or time limit), not that the product failed; the CI client exits 2 for it.","Pin the issuer's key (GET /attest/signing-key) when you verify: someone holding a different key can build a self-consistent certificate, and only pinning tells them apart."]},"buyer_facts":[{"label":"What decides pass or fail","value":"Only the assertions in your spec, judged in code on the live page. The model picks actions and never sees them."},{"label":"Typical run cost","value":"A handful of model actions, a fraction of a cent at the gateway list price (measured over many runs)."},{"label":"Data retention","value":"Hosted: certificates and screenshots-as-thumbnails for 24 hours, then deleted. Self-host: as long as you set."},{"label":"What leaves the box (hosted)","value":"Each step's element table and page text go to the gateway model; screenshots stay on the runner as hashes and 320 px thumbnails in the certificate. Secrets are replaced by {{NAME}} before anything is recorded or sent."},{"label":"Where it may run","value":"The fixture app, saucedemo.com, and https hosts whose domain your API key verified (DNS TXT or /.well-known file, valid 30 days). Self-host: any host you list."},{"label":"Spec format","value":"YAML or JSON: do or goto steps; assertions text_present, text_absent, url_contains, url_equals, title_contains, element_count, value_equals, value_contains."},{"label":"CI","value":"decosa_testrun.py (standard library): exit 0 pass, 1 fail, 2 error; writes and verifies the certificate; --repeat for a signed flake report."}],"data_handling":{"page":"/data#test-runs","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Hosted: certificates and screenshots-as-thumbnails for 24 hours, then deleted. Self-host: as long as you set.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"The site under test","route":"both","sends":"your-system","what":"The runner's browser loads the fixture app, saucedemo.com or https hosts your key verified (self-host: any host you list). Secrets are replaced by {{NAME}} before anything is recorded or sent to the model.","default":"always","off":null}]},"console":{"href":"/tools/developer/test-runs","input":"testruns","lanes":[{"id":"steps","title":"Test steps","kind":"list"},{"id":"verdict","title":"Verdict","kind":"markdown"},{"id":"certificate","title":"Run certificate","kind":"json"}],"samples":[{"n":1,"id":"kiln-cart-total@good","title":"Kiln cart total@good","deep_link":"/tools/developer/test-runs?sample=1&autorun=0"},{"n":2,"id":"kiln-cart-total@total-ignores-qty","title":"Kiln cart total@total ignores qty","deep_link":"/tools/developer/test-runs?sample=2&autorun=0"},{"n":3,"id":"kiln-cart-total@redesign","title":"Kiln cart total@redesign","deep_link":"/tools/developer/test-runs?sample=3&autorun=0"},{"n":4,"id":"sauce-checkout@standard_user","title":"Sauce checkout@standard user","deep_link":"/tools/developer/test-runs?sample=4&autorun=0"},{"n":5,"id":"kiln-signup@signup-broken","title":"Kiln signup@signup broken","deep_link":"/tools/developer/test-runs?sample=5&autorun=0"},{"n":6,"id":"kiln-promo@promo-wrong","title":"Kiln promo@promo wrong","deep_link":"/tools/developer/test-runs?sample=6&autorun=0"},{"n":7,"id":"kiln-remove@remove-wrong","title":"Kiln remove@remove wrong","deep_link":"/tools/developer/test-runs?sample=7&autorun=0"},{"n":8,"id":"kiln-checkout@checkout-disabled","title":"Kiln checkout@checkout disabled","deep_link":"/tools/developer/test-runs?sample=8&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/test-runs-hosted.md","selfhost":"/prompts/test-runs-selfhost.md","assemble":"/prompts/test-runs-assemble.md","mac":"/prompts/test-runs-mac.md"},"rehearsal":{"bundle":"/samples/test-runs.zip","bundle_url":"https://decosa.ai/samples/test-runs.zip","folder":"/samples/test-runs/","expected":"/samples/test-runs/expected.json","files":["/samples/test-runs/expected.json","/samples/test-runs/inputs/spec.yaml"],"bytes":1611,"checks":["the spec reads as valid, with 2 steps","the good build passes","every assertion on the good build passes","the planted total-ignores-qty build fails","it fails at the subtotal: $18.00 shown instead of $36.00","the good run's certificate verifies","a certificate with one assertion flipped no longer verifies","every model action in the certificate has a signed gateway receipt"],"licence":"Synthetic: Kiln & Co is a fictional fixture app shipped with decosa-api (AGPL-3.0-or-later); the spec was written for Decosa. No real shop or people.","about":"A two-step test spec for Kiln & Co, a fictional shop served from files on the server (nothing leaves it): the model agent opens a product page, picks quantity 2 and adds it to the cart, then the cart subtotal is asserted. It runs against the good build (must pass) and against the planted total-ignores-qty build (must fail at the subtotal). The good run's signed certificate must verify, and a copy with one assertion flipped must not.","run":{"containers":"docker compose exec api python scripts/rehearse.py test-runs","checkout":"python scripts/rehearse.py test-runs --bundle test-runs.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py test-runs"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=test-runs","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"alternate-vision-actions","gpu_gb":24.5,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/test-runs","json":"/use-cases/test-runs.json","metrics":"/metrics/test-runs","console":"/tools/developer/test-runs","console_sample":"/tools/developer/test-runs?sample=1&autorun=0","stack":"/tools/developer/test-runs#stack","try_live":"/tools/developer/test-runs","watch":"/tools/developer/test-runs","build":"/tools/developer/test-runs#build","self_host":"/tools/developer/test-runs#self-host","prompts":{"hosted":"/prompts/test-runs-hosted.md","selfhost":"/prompts/test-runs-selfhost.md","assemble":"/prompts/test-runs-assemble.md","mac":"/prompts/test-runs-mac.md"}}}