Skip to content
decosa
PreviewHostedSelf-hostSelf-host first for real data

Capture audit evidence from your admin screens

Each named screen captured with URL, time and user stamped, every setting checked against the baseline you approved, changes first, in a signed pack.

Held-out test33 / 33Screens given the right verdict on held-out consoles (held-out test)
On production7.5 smedian on production (2026-09-29); slower when the service is busy
List price~$0.037 per 10 screens set upmeasured, at list price

Built on: Agent flight recorder, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on The quarterly run needs only a CPU (headless Chromium). Setup and repairs call Qwen3.8-27B: 1x RTX PRO 6000 (96 GB) self-hosted, or the hosted gateway.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the capture audit evidence from your admin screens API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
evidence-runner

Use the hosted API

# Decosa evidence runner: try it on the hosted API, then verify packs (hosted API)

You are helping me evaluate the Decosa evidence runner. On the hosted API it runs only on made-up admin consoles (real
consoles sit behind our SSO and device checks, so real runs happen on our side: see the self-host prompt). Use only what
is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- What it does, the rails, the samples and limits: `GET https://api.decosa.ai/evidence/info` and `GET https://api.decosa.ai/evidence/samples`.

## Auth
`POST https://api.decosa.ai/demo/session` with `{"vertical": "evidence-runner"}` returns `{"token", ...}`; or an API key (`dk_…`)
enabled for `evidence-runner`, kept in `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer <token or key>`.

## A quarterly run on a made-up console (no model calls)
```bash
curl -sN https://api.decosa.ai/evidence/runs -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' -H 'Accept: text/event-stream' \
  -d '{"sample": "keystone", "mode": "certify", "quarter": "q3"}'
```
Events: `ready`, `person` (the sign-in hand-off), `approval` (who approved the baseline), one `capture` per screen (pass or
changed, baseline and current values, other settings on the screen that changed, URL, UTC time, PNG SHA-256), then
`certificate` and `done` (verdict, seconds, writes held, export links). Expect four screens changed since the baseline.
Download the pack: `GET https://api.decosa.ai/evidence/runs/<run_id>/export?format=zip` (same token).

## A setup run (the agent finds screens, read-only)
`{"sample": "keystone", "mode": "setup", "screens": ["2sv", "roles", "audit"]}` streams `screen` and `baseline` events with the
values code read, how each was read, model calls and seconds per screen, and a `receipt` event per model call.

## Verify a pack
`POST https://api.decosa.ai/testruns/verify` with `{"certificate": <certificate.json>}` → `ok: true`; change one byte and it fails.
Offline: `python -m decosa_api.verticals.evidence.cli verify --pack pack.zip` (also checks every PNG against the certificate).

Tell me what changed, how long the run took, and whether the certificate verified.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa evidence runner: run it on our side (local helper or in our network)

You are setting up the evidence runner for our own admin consoles and our own product's admin screens. It runs on our
side: a local helper attaches to the browser I am already signed in to and opens one tab of its own, or it runs inside
our network. Read-only: its tab blocks every write and it refuses toggles, pickers and save-like clicks. Follow the
assemble prompt on the tool page for installation; this prompt is the day-to-day workflow. Stop and ask me whenever
a step needs a decision.

1. Screens. Write `screens.json` with me: one entry per settings screen we screenshot today, with its SOC 2 or CMMC
   references and the settings to read on it. Prefer a console's own read-only API (Microsoft Graph, Google's Policy
   API, the AWS credential report, Okta's policies API) where it covers a setting; use this tool for the rest and for our
   own product's admin screens. For Microsoft 365 and Entra admin centers use `witness` only (I click, it stamps).
2. Setup (once per console, needs the model): `python -m decosa_api.verticals.evidence.cli setup --cdp http://127.0.0.1:<port> --origin https://<console> --screens screens.json --user <me> --out spec.json`.
   Show me each screen's `status` and `readings`. Screens `not_found` or flagged (text aimed at AI agents) stay with me.
3. Approve: `python -m decosa_api.verticals.evidence.cli approve --spec spec.json --by "<my name>" [--policy policy.json] --out approved.json`.
   A policy per setting is optional (`"on"`, `"off"`, `">=12"`, `"<=30"`, `"=text"`, `"keep"`).
4. Every quarter (no model): `python -m decosa_api.verticals.evidence.cli certify --cdp http://127.0.0.1:<port> --spec approved.json --user <me> --out evidence-<quarter>.zip`.
   Exit 0 = every screen matches; exit 1 = a change (the manifest lists what, which way, and other settings on the same
   screen that moved) or a screen not reached (the console changed: run setup for it and approve again).
5. Verify before filing: `python -m decosa_api.verticals.evidence.cli verify --pack evidence-<quarter>.zip`.
6. Keep the signing key (`~/.decosa/evidence-signing.pem`) safe and give the auditor its public key.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/evidence-runner.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py evidence-runner` (the api image carries the same bundle under /app/rehearsal/evidence-runner/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py evidence-runner --bundle evidence-runner.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the hosted service offers the made-up consoles only", "nine screens are captured", "exactly the four planted screens changed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"evidence-runner"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py evidence-runner

Download the mock-data bundle (1 KB, 10 checks)expected.json

Keystone Admin is a made-up identity admin console served inside the runner's own browser (nothing leaves it). The approved baseline was made by a setup run on last quarter's settings; this quarter four named settings changed and one setting next to a named one flipped. The run opens every screen from its saved route with no model calls, re-reads each setting, captures every screen, and signs a certificate. It must report exactly the planted changes, reach every screen, let no write through, and the certificate must verify while a copy with one check flipped must not.

What the rehearsal checks
  • the hosted service offers the made-up consoles only
  • nine screens are captured
  • exactly the four planted screens changed
  • the minimum password length change is found (12 to 14)
  • the setting that flipped next to a named one is reported
  • every saved route reached its screen
  • the quarterly run makes no model calls
  • no write reached the console
  • the certificate verifies
  • a certificate with one check flipped no longer verifies

Licence: Synthetic: the console, its company and its settings were written for Decosa (decosa-api, AGPL-3.0-or-later). No real product, tenant or person.

Prompt for your coding agent

# Decosa evidence runner: run it on our side (local helper or in our network)

You are setting up the evidence runner for our own admin consoles and our own product's admin screens. It runs on our
side: a local helper attaches to the browser I am already signed in to and opens one tab of its own, or it runs inside
our network. Read-only: its tab blocks every write and it refuses toggles, pickers and save-like clicks. Follow the
assemble prompt on the tool page for installation; this prompt is the day-to-day workflow. Stop and ask me whenever
a step needs a decision.

1. Screens. Write `screens.json` with me: one entry per settings screen we screenshot today, with its SOC 2 or CMMC
   references and the settings to read on it. Prefer a console's own read-only API (Microsoft Graph, Google's Policy
   API, the AWS credential report, Okta's policies API) where it covers a setting; use this tool for the rest and for our
   own product's admin screens. For Microsoft 365 and Entra admin centers use `witness` only (I click, it stamps).
2. Setup (once per console, needs the model): `python -m decosa_api.verticals.evidence.cli setup --cdp http://127.0.0.1:<port> --origin https://<console> --screens screens.json --user <me> --out spec.json`.
   Show me each screen's `status` and `readings`. Screens `not_found` or flagged (text aimed at AI agents) stay with me.
3. Approve: `python -m decosa_api.verticals.evidence.cli approve --spec spec.json --by "<my name>" [--policy policy.json] --out approved.json`.
   A policy per setting is optional (`"on"`, `"off"`, `">=12"`, `"<=30"`, `"=text"`, `"keep"`).
4. Every quarter (no model): `python -m decosa_api.verticals.evidence.cli certify --cdp http://127.0.0.1:<port> --spec approved.json --user <me> --out evidence-<quarter>.zip`.
   Exit 0 = every screen matches; exit 1 = a change (the manifest lists what, which way, and other settings on the same
   screen that moved) or a screen not reached (the console changed: run setup for it and approve again).
5. Verify before filing: `python -m decosa_api.verticals.evidence.cli verify --pack evidence-<quarter>.zip`.
6. Keep the signing key (`~/.decosa/evidence-signing.pem`) safe and give the auditor its public key.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/evidence-runner.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py evidence-runner` (the api image carries the same bundle under /app/rehearsal/evidence-runner/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py evidence-runner --bundle evidence-runner.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the hosted service offers the made-up consoles only", "nine screens are captured", "exactly the four planted screens changed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsCapture audit evidence from your admin screens on GeForce RTX 5090: use the Standard · setup and repairs with Qwen3.8-27B tier

The standard tier fits with changes: Qwen3.8-27B: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · setup and repairs with Qwen3.8-27B: what changesuses estimates

  • Qwen3.8-27B: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Browser session with the read-only gate, save...: decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Setup and repairs only: Qwen3.8-27B. ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 96 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Capture audit evidence from your admin screens, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Capture audit evidence from your admin screens on my hardware

Fetch https://decosa.ai/prompts/evidence-runner-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=evidence-runner)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · setup and repairs with Qwen3.8-27B (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Browser session with the read-only gate, save...: decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27), CPU
- Setup and repairs only: Qwen3.8-27B (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/evidence-runner-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 7.5 s · p95 10 s (3 runs) · ~$0 per run · 0 receipts

Loading the nightly status…

Self-host: verified 29 Sep 2026 · fresh clone, venv, tests, then the CLI attached over CDP to a separate Chromium profile standing in for the admin's own browser, against a made-up console served over plain HTTP

Measured cost to run: about $0.037 per 10 screens set up (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Setup 3 of 3 screens right in 46 s on the local model server; next quarter 2 changes found in 2 s (exit 1), 0 writes reached the console, the stand-in browser's own tab left untouched.

Known limits (6)
  • Made-up consoles only so far; no real tenant has been run.
  • Only the settings you name are asserted in the certificate; other settings on the same screen are compared and reported, not asserted.
  • The URL and time band is drawn onto each PNG by the runner (the certificate binds the file to its capture); it is not the browser's own address bar or the system clock.
  • Microsoft 365 and Entra admin centers: witnessed capture only (no agent navigation); use Microsoft Graph for settings.
  • Setup needs a person to approve the baseline values and to handle screens left for a person (pages with text aimed at AI agents, screens the agent could not find).
  • Hosted runs are demos on made-up consoles; real consoles run on your side.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on The quarterly run needs only a CPU (headless Chromium). Setup and repairs call Qwen3.8-27B: 1x RTX PRO 6000 (96 GB) self-hosted, or the hosted gateway.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the capture audit evidence from your admin screens API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

Named admin settings screens captured every quarter with URL, time and user stamped, each setting re-read against the baseline you approved, in a signed certificate. Read-only.

For the compliance lead or consultant who screenshots admin settings screens every quarter for SOC 2 or CMMC. Name the screens once; an agent finds each one, read-only, and code reads the settings you named; you approve those values as the baseline. Every quarter after that it runs with no model: saved routes reopen each screen, code re-reads every setting, each screen is captured on a fresh load and stamped, and changes since the baseline come first (with their direction, and any other setting on the same screen that moved). It runs on your side, in your own signed-in browser or inside your network, and its browser session blocks every write.

Deployment
Self-host first
Regulatory
Read 28 Sep 2026 (public terms only; not legal advice): Google Workspace (Terms, last modified 2 Sep 2026; Acceptable Use Policy, 13 Oct 2025), the AWS Customer Agreement (14 Aug 2026) and Acceptable Use Policy (1 Jul 2021), and the Okta Master Subscription Agreement (Q1 FY26) say nothing against a customer automating read-only views of its own tenant, and each offers read-only APIs (Cloud Identity Policy API, the IAM credential report and a SecurityAudit role, okta.policies.read) that should come first. Microsoft's Product Terms bar using an Online Service "to scrape or use other data extraction methods", so for Microsoft 365 and Entra this tool offers only witnessed capture (a person clicks, it stamps and signs) and Microsoft Graph is the route for settings. Okta's MSA also bars using the service for others as a managed service: run the tool as the customer's own. Auditors ask for a URL and date stamp on screenshots (Secureframe, 16 Jul 2026) and for the completeness and accuracy of system-generated evidence (PCAOB AS 1105.10). This tool captures and compares; it does not say a control is effective, and it is not an assessment.
Architecture
Text description

On the customer's side, a local helper attaches to the admin's own signed-in browser and opens one tab of its own. A read-only network gate blocks every write from that tab. Setup, once per console: the agent (Qwen3.8-27B) finds each named screen, code reads the settings, and a person approves the baseline. Every quarter: no model; saved routes reopen each screen, code re-reads every setting, a stamped screenshot is captured on a fresh load, and a signed certificate plus an evidence pack come out.

Architecture

At a glance

Data retention
Hosted demo: runs on made-up consoles are kept one hour for the session that made them, then deleted. Your own runs stay on your machine: the spec, the evidence packs and the signing key live where you run the helper.
What leaves the box
Quarterly runs: nothing (no model calls). Setup and repairs: screenshots and the element table of the admin screens the agent opens go to the model you choose (your own Qwen3.8-27B, or the hosted gateway with a key). Passwords never do: you sign in yourself.
Read-only by construction
The runner's tab blocks every write at the network level and no approval can release one; it refuses toggles, pickers and save-like clicks; every capture is a fresh load of the page.
Where it runs
On your side: a local helper attached to the browser you are already signed in to (it opens and gates one tab of its own), or inside your network. Your SSO, conditional access and device checks stay yours.
Typical quarter
A quarterly run takes seconds with no model calls; setup takes under a minute and a fraction of a cent per screen, once per console.
What it will not do
Say a control is effective, certify, or change a setting. It captures, compares with the baseline you approved, and signs.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    quarterly runs only, no GPU

    Replays an approved baseline spec: saved routes, every setting re-read in code, stamped captures, certificate, evidence pack, offline verify. No model; setup must come from the standard tier or a colleague's approved spec.

    Models
    • decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)
    Hardware
    Any CPU
    Quality evidence
    • Quarterly verdicts right, six made-up consoles52 / 52 screens (33 / 33 on held-out consoles); 16 / 16 changed screens caught; 0 false alarmsdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026
    Latency
    measured: about a second per screen
    Verification
    No proof yetNo model calls; the certificate is signed by the instance key.
  • In the hosted demo

    Standard

    setup and repairs with Qwen3.8-27B

    The agent finds each named screen read-only (two runs must agree) and repairs saved routes after a redesign; code reads the values; a person approves the baseline.

    Models
    • decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)
    • Qwen3.8-27B
    Hardware
    1x RTX PRO 6000 96 GB (measured) or the hosted gateway
    Quality evidence
    • Setup, held-out consoles after fixes25 / 26 screens found with every named setting read right; 0 wrong valuesdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026 (hosted gateway run)
    • Setup, first look at a console added after every fix8 / 10 (both misses from one bug, fixed after)decosa-api docs/evals/evidence-runner.md, 29 Sep 2026
    • Redesigned console, routes repaired19 / 19 screensdecosa-api docs/evals/evidence-runner.md, 29 Sep 2026
    Latency
    measured: under a minute per screen under shared load; a fraction of a cent per screen at list price
    Verification
    Setup decisions on the hosted gateway carry signed receipts.

Also runs on

  • Official read-only API exports (Graph, Policy API, AWS, Okta)decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)not builtThe first route where an API covers a setting; not built into this tool yet (compliance platforms already do it). Hardware: Any CPU.

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)
0 GBProof: partial
Setup and repairs only: finds each named screen (read-only) and points at rows when code cannot place a settingQwen3.8-27Bnvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27B · 96 GBNo proof yet

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api (local helper or in-network service):8445
    ${DECOSA_REGISTRY}/decosa-api:<tag> (publishing soon; build from source)

    The runner, the CLI (setup, approve, certify, verify, witness) and /testruns/verify.

  • Qwen3.8-27B on vLLM (optional):8114
    vllm/vllm-openai:v0.29.0

    Setup and repairs, self-hosted; or use the hosted gateway with a key.

Hardware

  • Any laptop or small VM (quarterly runs) Fits

    Headless Chromium on CPU; no GPU and no model calls.

  • 1x RTX PRO 6000 96 GB (setup, self-hosted) Fits

    Qwen3.8-27B NVFP4; or the hosted gateway instead.

Latency per lane

  • quarterly run, 9-10 screens, no model7.5 s

    Measuredmeasured on our server 2026-09-29: 6.3-10.2 s per run over 6 made-up consoles (0.7-1.5 s per screen)

  • setup, per screen (two agent runs must agree)25.0 s

    Measuredmeasured on our server 2026-09-29 through the hosted gateway under shared load: per-screen p50 17-43 s across four consoles

Notes

  • Measured only on made-up consoles that copy the shapes of real admin screens (identity admin, cloud console, a company's own product admin in iframes, an endpoint manager with dialogs, code-hosting organisation settings, an HR and payroll system). No real tenant yet.
  • Setup found the screen and read every named setting right on 19 of 20 development screens; on the first look at two held-out consoles 10 of 17 (four bugs found and fixed); after the fixes 25 of 26 on three held-out consoles; 8 of 10 on a sixth console run once at the end (both misses from one bug, since fixed). No wrong value was read in any run: what code cannot place is left for the person.
  • Quarterly runs gave the right verdict on every screen of all six consoles (52 of 52; 16 of 16 changed screens caught, no false alarms), with no model calls. A redesigned console: 19 of 19 screens handled (broken routes repaired by the agent, new values right, sent for approval).
  • No run changed a setting: in every run no write reached the console. A page carrying text aimed at AI agents stops the agent before the model reads it; that screen is captured by code only or left for the person.
  • Engine changes held MiniWoB held-out (181 of 252, main 180) and Gitea held-out (43 of 48, main 42).
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

evidence-runner/assemble-prompt.md97 lines
# Assemble the Decosa evidence runner on our side

You are setting up the evidence runner: once per admin console an agent finds each named settings screen (read-only)
and code reads the settings we name; we approve those values as the baseline. Every quarter after that it runs with no
model: it opens each screen from its saved route, re-reads every setting, captures a stamped screenshot on a fresh load
and signs a certificate. It must run on our side, because our consoles sit behind single sign-on, conditional access,
device checks and IP allow-lists. Work step by step, show me each command before you run anything with `sudo`, and stop
to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/evidence-runner.zip (1 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py evidence-runner` (the api image carries the same bundle under /app/rehearsal/evidence-runner/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py evidence-runner --bundle evidence-runner.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the hosted service offers the made-up consoles only", "nine screens are captured", "exactly the four planted screens changed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Runner: decosa-api (AGPL-3.0-or-later), CPU only. Browser automation: Playwright (Apache-2.0) with Chromium
  (BSD-3-Clause). Images: Pillow (MIT-CMU). Model for setup and repairs only: Qwen3.8-27B (Apache-2.0) on vLLM, or the
  hosted Decosa gateway with a `dk_` key (one receipt per call).
- Read-only by construction: the runner's browser tab blocks every write at the network level (no approval can release
  one), refuses toggles, pickers and save-like clicks, and captures each screen on a fresh load.
- Use a view-only admin role for the account the runner works under. Never give the runner a password: I sign in myself.
- Where a console's own read-only API covers a setting (Google Workspace Policy API, Microsoft Graph with
  Policy.Read.All, AWS SecurityAudit role, Okta `okta.policies.read`), prefer the API. For Microsoft 365 / Entra admin
  centers use witnessed capture (step 5) instead of agent navigation: Microsoft's Product Terms bar "scrape or ...
  other data extraction" (read 28 Sep 2026). Our own product's admin screens are fine for the agent.

## 1. Check the machine
1. Python 3.12 and `git`. For setup and repairs, either a GPU box running Qwen3.8-27B (one 96 GB card, or a 32 GB card
   with the FP8 weights) or a Decosa API key for the hosted gateway. Quarterly runs need no GPU.
2. Google Chrome (or Chromium) on the machine where I sign in to the consoles.
3. Disk: about 400 MB for the runner and Chromium, and 1-3 MB per quarterly evidence pack.

## 2. Install the runner (local helper)
```bash
git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>   # access required; a release with decosa_api/verticals/evidence/
cd decosa-api && python3 -m venv .venv && .venv/bin/pip install ".[testruns]"
.venv/bin/python -m playwright install chromium chromium-headless-shell
```
Model access for setup (pick one, in `~/.decosa/evidence.env`, mode 0600):
- hosted: `DECOSA_LLM_ROUTE=gateway`, `DECOSA_GATEWAY_URL=<gateway url>`, `DECOSA_GATEWAY_API_KEY=<dk_ key>`;
- our own GPU: `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://127.0.0.1:8114/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`
  (vLLM: `vllm/vllm-openai:v0.29.0 --model nvidia/Qwen3.8-27B-NVFP4 --served-model-name qwen3.8-27b`).

## 3. Attach to the browser I already use
1. Turn on remote debugging in my normal Chrome profile: open `chrome://inspect/#remote-debugging` and tick the box.
   Chrome writes the port to the `DevToolsActivePort` file in the profile directory (macOS:
   `~/Library/Application Support/Google/Chrome/DevToolsActivePort`); its first line is the port. Chrome asks me to
   Allow each new connection: I click Allow. (Recent Chrome versions ignore `--remote-debugging-port` for the default
   profile, so don't use that flag on it.) The port listens on 127.0.0.1 only; never forward it.
2. I sign in to the console in a normal tab (SSO, MFA, device check as usual).
3. The helper opens ONE tab of its own, gates only that tab and closes it at the end. It never reads my other tabs.
   If the console signs me out mid-run, the helper stops and asks me to sign in again in its tab.

## 4. Set up once, then run each quarter
1. Write `screens.json` with me, one entry per screen we screenshot today:
   `[{"id": "mfa", "name": "2-step verification settings", "controls": ["SOC 2 CC6.1"], "settings": ["Enforcement", "New user enrollment period"]}]`.
2. Setup (add `--hosts <sso host>` if the console redirects through a sign-in host): `set -a; . ~/.decosa/evidence.env; set +a; .venv/bin/python -m decosa_api.verticals.evidence.cli setup --cdp http://127.0.0.1:<port> --origin https://<console host> --screens screens.json --user <my admin email> --out spec.json`.
   Show me every screen's `status` and `readings` in `spec.json`. I approve the values (or fix the screen list and run
   setup again). Screens marked `not_found` or flagged (page text aimed at AI agents) stay with me.
3. Every quarter: `.venv/bin/python -m decosa_api.verticals.evidence.cli certify --cdp http://127.0.0.1:<port> --spec spec.json --user <my admin email> --out evidence-<quarter>.zip`.
   Exit 0 = every screen matches the baseline; exit 1 = something changed (the manifest lists what) or a screen was not
   reached (the console changed: run setup again for that screen and approve the new baseline).
4. Unzip the pack and check `manifest.csv`; `certificate.json` verifies with `POST /testruns/verify` on any Decosa API
   (or our own, step 6), and the CMMC evidence map reads it as a test-run certificate.

## 5. Witnessed capture (consoles whose terms bar agents)
`.venv/bin/python -m decosa_api.verticals.evidence.cli witness --cdp http://127.0.0.1:<port> --origin https://<console host> --user <me> --out witnessed.zip`:
I click to each screen myself and press Enter; it captures that tab, stamps URL, UTC time and "witnessed", and signs.

## 6. Optional: inside the network as a service
Run the decosa-api image (`docker build -f docker/api/Dockerfile --build-arg WITH_BROWSER=1 -t decosa-api:local .`) on a
host inside our network with `DECOSA_HOST=127.0.0.1`, a named data volume, and the model settings above; the same routes
(`/evidence/info`, `/testruns/verify`) are then available to our team. Bind to 127.0.0.1 or put it behind our SSO proxy.

## 7. Smoke test
1. `.venv/bin/python -m pytest -q tests/test_evidence.py tests/test_cu_readonly.py tests/test_cuharness.py`: all pass
   (made-up consoles, no network, no model).
2. Tell me the setup time per screen, the quarterly run time, and the certificate check summary you saw.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Capture SOC 2 evidence screenshots from admin consoles without the quarterly clicking: name the screens once, approve the values, and every quarter it opens each screen, re-reads each setting, stamps the capture and lists what changed.
Who it's for
Compliance leads at SaaS companies and the consultants and MSPs who prepare SOC 2 and CMMC evidence.
Where it runs
On your side: a local helper attached to your own signed-in browser, or self-hosted in your network. The hosted demo runs on made-up consoles only
Key numbers

On four held-out made-up consoles, quarterly runs gave the right verdict on 33 of 33 screens with no model calls, and setup read no wrong value in any run. No real tenant yet.

  • 33 / 33 Quarterly verdicts right on held-out consoles (no model) (test split, n = 33)
  • 25 / 26 Setup: screen found and every named setting read right, held-out after fixes (test split, n = 26)
  • 10 / 17 Setup, first look at two held-out consoles (test split, n = 17)
  • 7.5 s Median end-to-end run, hosted (QA sweep 2026-09-29)
All results, datasets and caveats
Models
Qwen3.8-27B finds each screen once (setup) and repairs routes when a console is redesigned; reading the settings, the quarterly replay, the stamps and the certificate are plain code
Where
On your side: a local helper attached to your own signed-in browser, or self-hosted in your network. The hosted demo runs on made-up consoles only
Checks
A decosa.test-run.v1 certificate per run (every setting re-read by code, stamped screenshots hashed in, signed); receipt per model call at setup
Output
Signed record or verdict · Structured data
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Does it take SOC 2 evidence screenshots for me?

Yes, for the admin screens you name. After a one-time setup it opens each screen from a saved route every quarter, re-reads the settings you named, captures the screen stamped with URL, time and the signed-in user, and lists what changed since the baseline you approved. It does not replace your compliance platform's API integrations.

Can it change my settings?

No. Its browser tab blocks every write at the network level and no approval can release one; it also refuses toggles, pickers and save-like clicks. In every measured run, no write reached a console.

Where does it run?

On your side: a local helper attaches to the browser you are already signed in to and opens one tab of its own, or you run it inside your network. Admin consoles sit behind single sign-on and device checks, so a browser on our servers could not reach them anyway.

Does it use AI every quarter?

No. The model is used only at setup (finding each screen once) and to repair a route after a console redesign. Quarterly runs are plain code and make no model calls.

What about Microsoft 365 and Entra?

Microsoft's Product Terms bar scraping and other data extraction, so there the tool offers witnessed capture only: you click to each screen and it stamps and signs. Use Microsoft Graph for the settings themselves.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Capture audit evidence from your admin screens

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.