Skip to content
decosa
LiveHostedSelf-hostMac

Run an end-to-end browser test

A browser run of your plain-language test, judged only by your assertions in code, ending in a signed certificate that fails the build on a failure.

Held-out test24/24Planted bugs and faulty test users caught (held-out test)
On production86 smedian on production (2026-09-25); slower when the service is busy
List price~$0.15 per 100 runsmeasured, at list price

Built on: Agent flight recorder, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the verified end-to-end test runs API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for the action model; the runner, browser and verifier run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
test-runs

Use the hosted API

# Decosa verified test runs: add certified end-to-end tests to this project (hosted API)

You are adding Decosa verified test runs to this project. A test is a YAML spec of plain-language steps, each with
assertions. Decosa's browser agent carries out the steps; pass or fail comes only from the assertions, checked in code
on the live page. Every run returns a signed run certificate to keep with the release. Use only what is listed below.
If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- What a run records and proves, the assertion kinds and the limits: `GET https://api.decosa.ai/testruns/info`.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, enabled for `test-runs`. Keep it in
   `DECOSA_API_KEY` (a CI secret), never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "test-runs"}` returns `{"token", ...}`. A demo
   session can test only the built-in targets (`fixture`, `saucedemo`), not our own site.

## Try it first (no setup)
```bash
curl -s https://api.decosa.ai/testruns/samples | python3 -c "import json,sys; print(json.load(sys.stdin)['samples'][2]['spec'])" > promo.yml
curl -sN https://api.decosa.ai/testruns/runs -H "Authorization: Bearer $DECOSA_API_KEY" -H 'Content-Type: application/json' \
  -d "$(python3 -c "import json; print(json.dumps({'spec': open('promo.yml').read(), 'target': 'fixture', 'build': 'promo-wrong'}))")" \
  | python3 -c "import json,sys; r=json.load(sys.stdin); print(r['summary']); json.dump(r['certificate'], open('certificate.json','w'))"
```
Expect `FAIL: 2 of 3 steps passed. step 3: [data-test=discount] is "-$2.40" (read "-$4.80")`: that build has a seeded
bug. With `"build": "good"` it passes. Then `POST https://api.decosa.ai/testruns/verify` with `{"certificate": ...}` returns
`ok: true`; change any value in the file and it returns `ok: false` naming the entry.

## Our own site: prove we own the domain (once)
1. `POST /testruns/domains` `{"domain": "staging.example.com"}` → a token and two ways to publish it.
2. Either a DNS TXT record `_decosa-challenge.<domain>` = the token (covers the domain and its subdomains), or serve the
   token as the body of `https://<host>/.well-known/decosa-verify.txt` (covers that exact host). Ask me which.
3. `POST /testruns/domains/check` `{"domain": "..."}` → `verified: true` for 30 days. Runs against an unverified host
   are refused with 403. Only public https hosts on port 443.

## Write the specs
Put them in `tests/e2e/*.yml`. Format:
```yaml
name: Checkout total
start: /login                      # path on the target
approve: []                        # buttons the agent may press although they look final (Pay, Finish, Delete ...)
steps:
  - do: Log in as 'qa@example.com' with password '{{TEST_PASSWORD}}'.
    expect:
      - url_contains: /dashboard
  - goto: /cart                    # plain navigation, no model
    expect:
      - element_count: {selector: "[data-test=cart-line]", equals: 1}
      - value_equals: {selector: "[data-test=total]", equals: "$23.00"}
```
- Every step needs `expect`. Kinds: `text_present`, `text_absent` (`{text, selector?, ignore_case?}` or a string),
  `url_contains`, `url_equals`, `title_contains`, `element_count` (`{selector, equals | at_least | at_most}`),
  `value_equals`, `value_contains` (`{selector, equals | contains}`). Prefer `data-test` selectors.
- A value the agent types must be in quotes in the step. Passwords: `'{{NAME}}'` in the spec, the value in `secrets`
  (the model types the placeholder, the browser fills in the secret; neither the model nor the certificate sees it).
- One job per step, 1-6 actions. A step ends when its assertions pass after an action, when the agent says done, or
  after `max_actions` (default 6). The first failed step ends the run; later steps are marked not run.
- Check a spec without running it: `POST /testruns/spec/check` `{"spec": "<yaml>"}` (no key needed).
- Use test accounts and test data only.

## Run it in CI
`curl -fsSO https://api.decosa.ai/testruns/sdk/decosa_testrun.py` (standard library only), then
`python3 decosa_testrun.py tests/e2e/checkout.yml --target https://staging.example.com --build-id "$GITHUB_SHA" --secret TEST_PASSWORD --out evidence/`
with `DECOSA_API=https://api.decosa.ai` and `DECOSA_API_KEY` set. Exit 0 = pass and verified, 1 = an assertion failed, 2 = error
or refused. Upload `evidence/` as a build artifact (keep it as long as your audit window). GitHub Actions:
```yaml
      - run: |
          curl -fsSO "$DECOSA_API/testruns/sdk/decosa_testrun.py"
          python3 decosa_testrun.py tests/e2e/checkout.yml --target https://staging.example.com --build-id "$GITHUB_SHA" --secret TEST_PASSWORD --out evidence/
        env: { DECOSA_API: "https://api.decosa.ai", DECOSA_API_KEY: "${{ secrets.DECOSA_API_KEY }}", TEST_PASSWORD: "${{ secrets.STAGING_TEST_PASSWORD }}" }
      - uses: actions/upload-artifact@v4
        if: always()
        with: { name: run-certificate, path: evidence/ }
```
`--repeat 3` runs it three times and writes a signed flake report (steps that both passed and failed).

## Endpoints
- `POST /testruns/runs` (token) `{spec, target?, build?, build_id?, secrets?, repeat?}` → with `Accept: text/event-stream`
  an SSE stream (ready, entry, step, receipt, certificate, flake, done); otherwise JSON when the run ends:
  `{verdict, run_id, summary, certificate, check}` (with `repeat` > 1: `certificates`, `flake`, `report`).
- `GET /testruns/runs/{id}` (the same token) → the certificate, for 24 hours.
- `POST /testruns/verify` `{certificate}` or `{report}` (no token) → `{ok, summary, checks, bad, test_steps, verdict}`.
- 429 with `Retry-After` when the browsers are busy; 402 when the key's daily model budget is used up.

## Rules
- Never treat the model's "done" as a pass; only the verdict counts, and only a certificate that verifies.
- Never test a site we do not own or have written permission to test.
- Show me the failing step, its expected value and the value read from the page when a run fails.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa verified test runs: run it yourself (containers)

You are setting up Decosa's verified test runs on this machine, so the staging site, test accounts, screenshots and
certificates stay here. The runner (decosa-api with a headless Chromium) needs no GPU; Qwen3.8-27B chooses the
actions. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", build decosa-api from source with
`docker build -f docker/api/Dockerfile --build-arg WITH_BROWSER=1 -t decosa-api:local .` (the browser is required),
and tell me. Do not substitute other images.

Hardware: any Linux x86_64 machine for the runner (2 GB RAM per concurrent browser); 1x RTX PRO 6000 96 GB (or a
32 GB card) for Qwen3.8-27B. Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/test-runs.zip (2 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py test-runs` (the api image carries the same bundle under /app/rehearsal/test-runs/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py test-runs --bundle test-runs.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the spec reads as valid, with 2 steps", "the good build passes", "every assertion on the good build passes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Rules you must keep
- Bind every port to 127.0.0.1. CI reaches it over a VPN or an authenticated proxy, never an open port.
- Only test sites I own or may test, with test accounts. List them in `DECOSA_TESTRUNS_TARGETS` (host names); add
  `DECOSA_TESTRUNS_ALLOW_PRIVATE=1` only if they are on a private network. Everything else is refused unless its domain
  is verified (DNS TXT or /.well-known file), exactly like the hosted service.
- The signing key is created on first start in the data volume (`attest/ed25519.pem`, 0600). Back it up; never print
  it; keep it away from the people whose releases the certificates vouch for.
- Passwords go in the run's `secrets`, never in a spec.

## Steps
1. Docker (and the NVIDIA Container Toolkit for the model): if `docker compose version` fails, install it from the
   official Docker instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `api` and `llm` services. The `api` service needs `shm_size: 1gb` for Chromium.
3. In `.env`: `DECOSA_LLM_ROUTE=direct`, `DECOSA_TESTRUNS_TARGETS=<your staging hosts>`, `DECOSA_TESTRUNS_MAX=2`,
   `DECOSA_TESTRUNS_TTL_S=86400`, `DECOSA_ADMIN_SECRET=<long random>`.
4. Start: `docker compose up -d`. Wait for `curl -fsS http://localhost:<PORT>/testruns/info` to show `"available": true`.
5. Mint a key: `POST /v1/keys` with the admin secret and `{"label": "ci", "verticals": ["test-runs"]}`.
6. Smoke test: `curl -fsSO localhost:<PORT>/testruns/sdk/decosa_testrun.py`; save sample 3 from `/testruns/samples`
   as `promo.yml`; run `DECOSA_API=http://127.0.0.1:<PORT> DECOSA_API_KEY=dk_... python3 decosa_testrun.py promo.yml --target fixture --build good --out ev/`
   (expect exit 0) and again with `--build promo-wrong` (expect exit 1, step 3 read "-$4.80"). Flip one `"pass"` in a
   certificate and POST it to `/testruns/verify`: it must fail and name the assertion.
7. Report back: health, the signing key id from `/attest/signing-key`, the two verdicts, and the run times you saw.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/test-runs-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"test-runs"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py test-runs

Download the mock-data bundle (2 KB, 8 checks)expected.json

A two-step test spec for Kiln & Co, a fictional shop served from files on the server (nothing leaves it): the model agent opens a product page, picks quantity 2 and adds it to the cart, then the cart subtotal is asserted. It runs against the good build (must pass) and against the planted total-ignores-qty build (must fail at the subtotal). The good run's signed certificate must verify, and a copy with one assertion flipped must not.

What the rehearsal checks
  • the spec reads as valid, with 2 steps
  • the good build passes
  • every assertion on the good build passes
  • the planted total-ignores-qty build fails
  • it fails at the subtotal: $18.00 shown instead of $36.00
  • the good run's certificate verifies
  • a certificate with one assertion flipped no longer verifies
  • every model action in the certificate has a signed gateway receipt

Licence: Synthetic: Kiln & Co is a fictional fixture app shipped with decosa-api (AGPL-3.0-or-later); the spec was written for Decosa. No real shop or people.

Prompt for your coding agent

# Decosa verified test runs: run it yourself (containers)

You are setting up Decosa's verified test runs on this machine, so the staging site, test accounts, screenshots and
certificates stay here. The runner (decosa-api with a headless Chromium) needs no GPU; Qwen3.8-27B chooses the
actions. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", build decosa-api from source with
`docker build -f docker/api/Dockerfile --build-arg WITH_BROWSER=1 -t decosa-api:local .` (the browser is required),
and tell me. Do not substitute other images.

Hardware: any Linux x86_64 machine for the runner (2 GB RAM per concurrent browser); 1x RTX PRO 6000 96 GB (or a
32 GB card) for Qwen3.8-27B. Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/test-runs.zip (2 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py test-runs` (the api image carries the same bundle under /app/rehearsal/test-runs/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py test-runs --bundle test-runs.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the spec reads as valid, with 2 steps", "the good build passes", "every assertion on the good build passes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Rules you must keep
- Bind every port to 127.0.0.1. CI reaches it over a VPN or an authenticated proxy, never an open port.
- Only test sites I own or may test, with test accounts. List them in `DECOSA_TESTRUNS_TARGETS` (host names); add
  `DECOSA_TESTRUNS_ALLOW_PRIVATE=1` only if they are on a private network. Everything else is refused unless its domain
  is verified (DNS TXT or /.well-known file), exactly like the hosted service.
- The signing key is created on first start in the data volume (`attest/ed25519.pem`, 0600). Back it up; never print
  it; keep it away from the people whose releases the certificates vouch for.
- Passwords go in the run's `secrets`, never in a spec.

## Steps
1. Docker (and the NVIDIA Container Toolkit for the model): if `docker compose version` fails, install it from the
   official Docker instructions.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Read it.
   Keep the `api` and `llm` services. The `api` service needs `shm_size: 1gb` for Chromium.
3. In `.env`: `DECOSA_LLM_ROUTE=direct`, `DECOSA_TESTRUNS_TARGETS=<your staging hosts>`, `DECOSA_TESTRUNS_MAX=2`,
   `DECOSA_TESTRUNS_TTL_S=86400`, `DECOSA_ADMIN_SECRET=<long random>`.
4. Start: `docker compose up -d`. Wait for `curl -fsS http://localhost:<PORT>/testruns/info` to show `"available": true`.
5. Mint a key: `POST /v1/keys` with the admin secret and `{"label": "ci", "verticals": ["test-runs"]}`.
6. Smoke test: `curl -fsSO localhost:<PORT>/testruns/sdk/decosa_testrun.py`; save sample 3 from `/testruns/samples`
   as `promo.yml`; run `DECOSA_API=http://127.0.0.1:<PORT> DECOSA_API_KEY=dk_... python3 decosa_testrun.py promo.yml --target fixture --build good --out ev/`
   (expect exit 0) and again with `--build promo-wrong` (expect exit 1, step 3 read "-$4.80"). Flip one `"pass"` in a
   certificate and POST it to `/testruns/verify`: it must fail and name the assertion.
7. Report back: health, the signing key id from `/attest/signing-key`, the two verdicts, and the run times you saw.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/test-runs-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsVerified end-to-end test runs on GeForce RTX 5090: use the Standard · agent steps, one 96 GB card (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1× RTX 5090 32 GB (fits): Qwen3.8-27B with a shorter context (about 700 prompt tokens per action is well inside it). Not measured here.

Standard · agent steps, one 96 GB card (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Runner: decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client. CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Action model: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'.
  • Headless browser that carries out the steps a...: Playwright 1.58 with Chromium headless shell. CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Verified end-to-end test runs, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Verified end-to-end test runs on my hardware

Fetch https://decosa.ai/prompts/test-runs-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=test-runs)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · agent steps, one 96 GB card (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Runner: decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client, CPU
- Action model: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Headless browser that carries out the steps a...: Playwright 1.58 with Chromium headless shell, CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/test-runs-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --browser

Mac prompt for your coding agent

# Decosa Verified end-to-end test runs: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Verified end-to-end test runs on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/test-runs.zip (2 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py test-runs` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the spec reads as valid, with 2 steps", "the good build passes", "every assertion on the good build passes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Action model: picks the next click, typing or selection from a numbered element table (never pass or fail) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Headless browser that carries out the steps and reads the assertions | CPU | Playwright's macOS arm64 Chromium | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh --browser`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py test-runs`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --browser --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 86 s · ~$0.002 per run · 6 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone, image built with WITH_BROWSER=1, compose up, sample against local model servers

Measured cost to run: about $0.15 per 100 runs (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Images build, the service starts, the sample passes on the good build (exit 0, 9 s) and fails on the seeded bug (exit 1, 19 s) against the already-running local Qwen3.8-27B vLLM (host network, no llm service started); a flipped assertion fails /testruns/verify; a private http target listed in DECOSA_TESTRUNS_TARGETS passes and an unlisted one is refused. Named volume for /data. Model-server startup itself not re-verified.

Known limits (5)
  • Eval cases are simple shop flows written for it (plus Sauce Labs' public site); expect more stuck steps on complex apps.
  • The action model reads the element table, not pixels: canvas-heavy and cross-origin iframe apps are not supported.
  • Assertions read the DOM; layout and visual regressions are out of scope.
  • A run takes about 1.5 minutes when the shared GPU is busy (seconds when it is quiet); a scripted Playwright test is faster and free.
  • Hosted runs only reach the fixture app, saucedemo.com and domains you verified; hosted certificates are kept 24 hours.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the verified end-to-end test runs API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for the action model; the runner, browser and verifier run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Plain-language browser tests judged only by assertions in code, each run sealed into a signed certificate you can keep as release evidence.

You write a test as YAML: plain-language steps ('Add one Linen Apron to the cart') and, for each step, assertions (text present, URL, element count, value equals). A headless browser carries out each step; Qwen3.8-27B picks the clicks and typing from the page's element table, with a signed receipt per action, and never sees the assertions. Pass or fail comes only from the assertions, read from the live page and judged in code, and the first failed step ends the run. Every action, every assertion's expected and actual value, and the verdict are chained and signed into a run certificate that names the spec, the target and the build; anyone can verify it and re-derive the verdict. A CLI runs specs in CI, fails the build on a failed assertion and keeps the certificate as change-management evidence. For small SaaS teams without QA staff, and teams that have to show an auditor that a release was tested.

Deployment
Hosted or self-host
Regulatory
Checked 2026-09-25. SOC 2 (AICPA Trust Services Criteria, CC8.1) and ISO/IEC 27001:2022 (Annex A 8.29 security testing in development and acceptance, 8.32 change management) expect evidence that changes were tested and approved before release; neither requires a particular tool, a signature or a hash chain. A run certificate is one piece of evidence an auditor can sample and re-check; it does not make a change process compliant, does not record who approved the release, and covers only what the spec asserts. Testing a website without the owner's permission can be unlawful (for example under the US Computer Fraud and Abuse Act or the UK Computer Misuse Act); the hosted service therefore only visits its own fixture app, Sauce Labs' public test site and domains whose owner proved control with a DNS TXT record or a /.well-known file. Use test accounts and test data: screenshots of a staging site can contain personal data, and under the GDPR minimisation and storage limits apply (hosted certificates are kept 24 hours; self-host keeps everything on your box). Passwords go in run secrets and never reach the model or the certificate. Model licence: Apache-2.0 (Qwen3.8-27B). Not legal advice.
Architecture
Text description

A test spec with plain-language steps and assertions comes from CI or the console, with a target (the fixture app, saucedemo.com or a domain proved by DNS TXT or a well-known file). The runner in decosa-api opens a headless Chromium in which only the target host resolves. For each step Qwen3.8-27B picks the next action from the element table, with a signed receipt from our gateway; guards run in code; the model never sees the assertions. The assertions are then read from the live page and judged in code, and the first failed step ends the run. Every action, assertion value and the verdict are hash-chained and signed into a run certificate naming the spec, target and build. Outputs: the certificate, verification in a browser or at POST /testruns/verify that re-derives the verdict, a CI exit code, and a signed flake report. The runner, browser, test secrets and signing key stay on your box when self-hosted.

Architecture

At a glance

What decides pass or fail
Only the assertions in your spec, judged in code on the live page. The model picks actions and never sees them.
Typical run cost
A handful of model actions, a fraction of a cent at the gateway list price (measured over many runs).
Data retention
Hosted: certificates and screenshots-as-thumbnails for 24 hours, then deleted. Self-host: as long as you set.
What leaves the box (hosted)
Each step's element table and page text go to the gateway model; screenshots stay on the runner as hashes and 320 px thumbnails in the certificate. Secrets are replaced by {{NAME}} before anything is recorded or sent.
Where it may run
The fixture app, saucedemo.com, and https hosts whose domain your API key verified (DNS TXT or /.well-known file, valid 30 days). Self-host: any host you list.
Spec format
YAML or JSON: do or goto steps; assertions text_present, text_absent, url_contains, url_equals, title_contains, element_count, value_equals, value_contains.
CI
decosa_testrun.py (standard library): exit 0 pass, 1 fail, 2 error; writes and verifies the certificate; --repeat for a signed flake report.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    navigation-only specs, any CPU

    Every step is a goto with assertions: no model, no GPU, still a signed certificate. For smoke checks of pages, not flows.

    Models
    • decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client
    • Playwright 1.58 with Chromium headless shell
    Hardware
    Any Linux machine with 2 GB RAM per browser
    Quality evidence
    • Navigation steps in the eval86 goto steps, all judged by the same assertion code as the agent stepsdecosa-api docs/evals/test-runs.md
    Latency
    estimate: a goto step and its assertions take under a second plus page load; no model wait
    Verification
    Proof: partialNo model calls, so no receipts; the certificate is attested by the box's own key.
  • In the hosted demo

    Standard

    agent steps, one 96 GB card (hosted demo)

    Qwen3.8-27B carries out the plain-language steps through the flight recorder's /decide, a receipt per action; assertions judge.

    Models
    • decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client
    • Qwen3.8-27B (NVIDIA NVFP4)
    • Playwright 1.58 with Chromium headless shell
    Hardware
    1× RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Verdict agrees with a hand-written Playwright test (52 cases: 5 specs x 8 fixture builds, 3 saucedemo specs x 4 users)52/52 first runs; 120/120 over all runsdecosa-api docs/evals/test-runs.md, 2026-09-25
    • Seeded bugs and faulty users caught24/24 runs, each at the same step as the scripted test; 0/96 false fails on working builds, including a UI redesigndecosa-api docs/evals/test-runs-results.json
    • Flaky cases over 2-4 repeats0/52decosa-api docs/evals/test-runs.md
    • Tampered certificates caught (11 alterations)1,320/1,320 with the issuer key pinneddecosa-api docs/evals/test-runs-results.json
    Latency
    measured: seconds per action and a minute or two per run with the shared GPU saturated; seconds for a short run on a quiet gateway
    Verification
    Proof: strongEvery action is gateway-receipted and bound to its output hash in the chain.

Also runs on

  • Vision actions for canvas and iframe appsHolo-3.1-35B-A3B (NVFP4)not servedTry the standard decider's own vision tower first (Qwen3.8-27B, same weights; screenshots took our MiniWoB++ run from 25.9% to 55.8%), then a dedicated grounder such as Holo-3.1 when pixel grounding is needed. Needs image-input receipts on the gateway. Hardware: 1x 96 GB card (estimate).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client
0 GBProof: partial
Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57.6 GBProof: strongIn the hosted demo
Headless browser that carries out the steps and reads the assertionsPlaywright 1.58 with Chromium headless shell
0 GBNo proof yetIn the hosted demo
Vision action model for canvas and iframe apps (controls not in the DOM)Holo-3.1-35B-A3B (NVFP4)Hcompany/Holo-3.1-35B-A3B-NVFP4 on Hugging Face (opens in a new tab)
35B (3B active) · about 24 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • decosa_testrun.py (CI client)Apache-2.0

    Served at GET /testruns/sdk/decosa_testrun.py. Runs a spec file, writes and verifies the certificate, exits 0 pass / 1 fail / 2 error; --repeat writes a signed flake report.

  • Kiln & Co (fictional fixture app)Apache-2.0

    Seven static pages served to the browser by request interception, in 8 builds: good, a redesign and 6 seeded bugs. Nothing is sold.

  • Built-in public target; its test users (standard_user, problem_user, locked_out_user, error_user) and password are printed on its login page.

  • scripts/testruns_eval.pyApache-2.0

    Agent verdicts against hand-written Playwright tests, repeats for flake rate, and the tamper set (decosa-api).

Services

  • decosa-api (test-run routes):8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    POST /testruns/runs (SSE or JSON), /testruns/verify, /testruns/spec/check, /testruns/domains and /domains/check; GET /testruns/info, /samples, /runs/{id}, /sdk/decosa_testrun.py. Build with WITH_BROWSER=1.

  • Action model (vLLM, or our gateway):8114
    vllm/vllm-openai:v0.29.0

    Qwen3.8-27B for the actions. Navigation-only specs need no model.

Hardware

  • Any CPU, no GPU Fits

    The runner and browser: about 2 GB RAM per concurrent browser. Specs made only of goto steps and assertions need nothing else.

  • 1× RTX PRO 6000 96 GB Fits

    Qwen3.8-27B NVFP4 for the actions; measured on our server.

  • 1× RTX 5090 32 GB Fits

    Qwen3.8-27B with a shorter context (about 700 prompt tokens per action is well inside it). Not measured here.

Latency per lane

  • one action (Qwen3.8-27B through the gateway)12.8 s

    Measuredmeasured on our server 2026-09-25: median of 692 actions, 10-90% 9.9-21.9 s, with the shared model server saturated by other evals

  • one action on a quiet gateway900 ms

    Measuredmeasured on our server 2026-09-25: 0.5-1.4 s per action on a 3-step run

  • whole run (3-4 test steps, about 6 actions)86.1 s

    Measuredmeasured on our server 2026-09-25: median of 120 eval runs under that load; 7.8 s on a quiet gateway, 9-19 s self-hosted on the direct route

  • verify a certificate10 ms

    Estimateestimate from the flight recorder's 5 ms for a 44-entry record; certificates here have 17-60 entries

Notes

  • The model chooses actions only. It is never shown the assertions, and a 'done' that the assertions contradict gets one generic 'not finished' note, never what is checked, so the agent cannot steer towards a pass.
  • A step ends when its assertions pass after an action (unless they already passed before the step), when the model says done, or after max_actions. The first failed step ends the run; later steps are marked not run.
  • Verdict error means the run could not be carried out (start page, model or time limit), not that the product failed; the CI client exits 2 for it.
  • Pin the issuer's key (GET /attest/signing-key) when you verify: someone holding a different key can build a self-consistent certificate, and only pinning tells them apart.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

test-runs/assemble-prompt.md131 lines
# Assemble Decosa verified test runs on this machine

You are setting up verified end-to-end test runs: YAML specs of plain-language steps with assertions, carried out by a
headless browser agent, judged only by the assertions (in code), each run sealed into a signed run certificate that
anyone can re-verify. Self-hosting keeps the staging site, test accounts and screenshots on this machine. Work step by
step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/test-runs.zip (2 KB, 8 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py test-runs` (the api image carries the same bundle under /app/rehearsal/test-runs/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py test-runs --bundle test-runs.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the spec reads as valid, with 2 steps", "the good build passes", "every assertion on the good build passes"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Runner: decosa-api (AGPL-3.0-or-later), CPU only. Browser: Playwright (Apache-2.0) with Chromium (BSD-3-Clause). Thumbnails:
  Pillow (MIT-CMU). YAML specs: PyYAML (MIT). Action model: Qwen3.8-27B (Apache-2.0) on vLLM.
- Bind every port to 127.0.0.1. Only point it at sites I own or may test, with test accounts and test data.
- The signing key is the trust anchor. It is created on first start in the data directory; back it up, never print it,
  and keep it away from the people whose releases the certificates vouch for.
- Be honest about what a certificate proves: which spec ran on which target and build, what each assertion read and
  that the verdict follows from those values, and that nothing changed after signing. Not that the site works for
  real users, and nothing the spec does not assert. The model chooses actions only; it never decides pass or fail.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for Qwen3.8-27B (an RTX PRO 6000 96 GB runs the NVFP4 weights with room to
   spare; on a card without NVFP4 use `Qwen/Qwen3.8-27B-FP8`). Driver 570 or newer. No GPU is needed for the runner.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 2.5 GB for the runner image with Chromium, about 25 GB for the model, and about 0.3-1 MB per certificate.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). Until then build it from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/testruns/`, and `docker build -f docker/api/Dockerfile --build-arg WITH_BROWSER=1 -t decosa-api:local .`
  (`WITH_BROWSER=1` is required: it installs the headless Chromium the runner drives).
- Without Docker: `python -m venv .venv && .venv/bin/pip install ".[testruns]" && .venv/bin/playwright install --with-deps chromium`
  in the checkout, then `.venv/bin/python -m decosa_api`.
- Model: `vllm/vllm-openai:v0.29.0` with weights `nvidia/Qwen3.8-27B-NVFP4`.

## 3. docker-compose.yml
Write this in `~/decosa/testruns/`. If a Qwen3.8-27B server already runs on this machine, drop `llm` and `depends_on`
and point `DECOSA_LLM_URL` at it.

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--language-model-only", "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/v1/models', timeout=4)"], interval: 30s, retries: 20 }
  api:
    image: decosa-api:local          # or ${DECOSA_REGISTRY}/decosa-api:<tag> once published
    ports: ["127.0.0.1:8445:8445"]
    shm_size: "1gb"                  # Chromium needs more than Docker's 64 MB default
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_TESTRUNS_MAX: "2"                  # browsers at once
      DECOSA_TESTRUNS_TTL_S: "86400"            # certificates kept on the box (export them to your archive)
      DECOSA_TESTRUNS_TARGETS: ${TEST_TARGETS}  # hosts this box may test without domain verification
      DECOSA_TESTRUNS_ALLOW_PRIVATE: "1"        # allow those hosts to be private addresses (a LAN staging box)
      DECOSA_FLIGHT_DEMO: "0"
      DECOSA_ADMIN_SECRET: ${DECOSA_ADMIN_SECRET}
    volumes: ["decosa-data:/data"]     # a named volume: the image runs as uid 10001, so a root-owned bind mount fails
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/testruns/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data: {}
```

In `~/decosa/testruns/.env` (mode 0600): `DECOSA_ADMIN_SECRET=<a long random string>` and
`TEST_TARGETS=staging.internal.example` (comma-separated host names; leave empty to allow only the built-in fixture app
and verified domains). A target listed here may use http and any port. On the first start the api service creates this
box's Ed25519 key in the `decosa-data` volume (`/data/attest/ed25519.pem`, 0600; back it up with
`docker compose cp api:/data/attest/ed25519.pem ./ed25519.pem.bak` and keep that copy safe). Actions on the direct route get receipts signed with that key
(status `attested`): an attestation by me, the operator, not a gateway receipt.

## 4. Smoke test
1. `docker compose up -d`, then `curl -s localhost:8445/testruns/info | python3 -m json.tool | head -40`: `available` must
   be `true` (the browser is installed) and `decide.route` `direct`.
2. Mint a key: `curl -s -XPOST localhost:8445/v1/keys -H "authorization: Bearer $DECOSA_ADMIN_SECRET" -H 'content-type: application/json' -d '{"label":"ci","verticals":["test-runs"]}'`.
   Then `export DECOSA_API_KEY=<the dk_ key>` in this shell (and keep it in your CI secrets).
3. Run the built-in sample on the good build and on a seeded bug:
   ```bash
   curl -s localhost:8445/testruns/samples | python3 -c "import json,sys; print(json.load(sys.stdin)['samples'][2]['spec'])" > promo.yml
   curl -fsSO localhost:8445/testruns/sdk/decosa_testrun.py
   DECOSA_API=http://127.0.0.1:8445 python3 decosa_testrun.py promo.yml --target fixture --build good --out ev/; echo "exit $?"
   DECOSA_API=http://127.0.0.1:8445 python3 decosa_testrun.py promo.yml --target fixture --build promo-wrong --out ev/; echo "exit $?"
   ```
   Expect `verdict PASS`, `verified: ...` and `exit 0`, then step 3 failing with `read "-$4.80"` and `exit 1`.
4. Tamper check: in one certificate in `ev/`, change a `"pass": false` to `true` and POST it to `/testruns/verify` as
   `{"certificate": ...}`: it must return `ok: false` and name the assertion.
5. Tell me the run times you saw and the signing key id from `GET /attest/signing-key`.

## 5. Use it
- Specs live in the repository (`tests/e2e/*.yml`); the format is at `GET /testruns/info` and in the hosted prompt on
  the site. Passwords go in `secrets` (`--secret NAME` in the CLI), never in the spec.
- CI: run `decosa_testrun.py` against this box (over a VPN or an authenticated proxy, never an open port) and keep
  `evidence/` with the release. Exit 1 on a failed assertion fails the build.
- The site: set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in `.env.local` to use the console against this box.
  Contract: `API_CONTRACT.md`, section "Verified end-to-end test runs".

Off by default. Joining serves other people's requests on this GPU; do not enable it on a box that holds test
credentials for private systems. If I ask for it, follow the provider guide at `/provide` on the site, and do not
enable it without my explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Plain-language browser tests judged only by assertions in code, each run sealed into a signed certificate you can keep as release evidence.
Who it's for
Teams in software and ai ops and compliance and trust.
Where it runs
Hosted (verified domains) or self-host
Key numbers
  • 52/52 Verdict agrees with the scripted Playwright test, first run of each case (test split, n = 52)
  • 120/120 Verdict agrees, all runs (test split, n = 120)
  • 24/24 Seeded bugs and faulty users caught (test split, n = 24)
  • 86.1 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B for actions; assertions in code
Where
Hosted (verified domains) or self-host
Checks
Receipted actions; signed run certificate whose verdict recomputes
Output
Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Who decides pass or fail in Decosa's verified end-to-end test runs?

Only the assertions in your spec. In verified end-to-end test runs, Qwen3.8-27B picks clicks and typing from the page's element table and never sees the assertions. Pass or fail comes from assertions such as text present, URL, element count and value equals, read from the live page and judged in code, and the first failed step ends the run. The model cannot mark a step as passed.

How accurate are the verified end-to-end test runs?

In the eval, verified end-to-end test runs agreed with hand-written Playwright tests in 120/120 runs, caught 24/24 seeded bugs and faulty users at the same step, had 0/96 false fails on working builds including a redesign, and caught 1,320/1,320 tampered certificates. Caveat: the fixture app, bugs and specs were written by the same author as the runner, and simple shop flows are not a large product.

Can a verified test run certificate serve as SOC 2 CC8.1 evidence?

A run certificate is one piece of evidence an auditor can sample and re-check. SOC 2 CC8.1 and ISO/IEC 27001:2022 Annex A 8.29 and 8.32 expect evidence that changes were tested before release, but neither requires a particular tool or signature. The certificate names the spec, target and build; it does not by itself satisfy a change process, does not record who approved the release, and covers only what the spec asserts.

How do verified end-to-end test runs fit into CI?

Verified end-to-end test runs ship a standard-library CI client, decosa_testrun.py, that exits 0 on pass, 1 on fail and 2 on error, writes and verifies the certificate, and accepts --repeat for a signed flake report. Specs are YAML or JSON with do or goto steps. A run costs about $0.0015 at the gateway list price, but a scripted Playwright test is faster and free.

What can't the verified end-to-end test runs test?

The action model in verified end-to-end test runs reads the element table, not pixels, so canvas-heavy and cross-origin iframe apps are not supported. Assertions read the DOM, so layout and visual regressions are out of scope. A run takes about 1.5 minutes when the shared GPU is busy and seconds when it is quiet. The eval covers simple shop flows; expect more stuck steps on complex apps.

Which sites can hosted verified test runs visit?

Hosted verified end-to-end test runs only reach the fixture app, saucedemo.com and https hosts whose domain your API key verified by DNS TXT record or a /.well-known file, valid 30 days, because testing a site without the owner's permission can be unlawful. Hosted certificates are kept 24 hours. Self-hosted, the runner visits any host you list and keeps data as long as you set.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Verified end-to-end test runs

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.