Skip to content
decosa
LiveHostedSelf-host

Draft VEX for scanner findings

A draft VEX statement for every scanner finding, with the evidence from the image itself, ready for a security engineer to sign off.

Held-out test154 / 159'Not affected' calls that match Canonical's own statements (held-out test)
On production14 smedian on production (2026-09-26); slower when the service is busy
List price~$0.12 per 100 findingsmeasured, at list price

Built on: Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the signed vex triage API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the collector, the checks, rules mode and the signing run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
vex-triage

Use the hosted API

# Decosa signed VEX triage: use the hosted API

You are wiring Decosa's VEX triage into this project (a CI pipeline, a release script or a security dashboard). It takes a
scanner report for one container image and returns draft VEX statements with their evidence, signed OpenVEX and CycloneDX
VEX documents, and a signed evidence record. Every model call has a signed receipt. Use only what is listed below; if you
need something else, stop and ask me.

Per finding the answer is one of:
- `not_affected`, with a CISA justification (`component_not_present`, `vulnerable_code_not_present`,
  `vulnerable_code_not_in_execute_path`; the other two are left to a person);
- `affected`, with an action;
- `fixed`;
- `under_investigation`.

`not_affected` is written only when a code check found positive evidence in the image or the advisory ranges. Everything
else stays under investigation.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The image never goes to the API.** You send the scanner report, optionally the SBOM, and the evidence bundle the
  collector writes on your machine (search hits, library loader chains, config excerpts with secret-looking lines
  redacted). Without the bundle nothing can be marked `not_affected`.
- This is a triage aid. Never publish the VEX without a security engineer's review, and never tell anyone an image is
  secure because of it.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "vex-triage"}` returns `{"token", "expires_at", "budget"}`.
   The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response). Over a limit you get HTTP 429
   with `Retry-After`. A demo token runs one triage at a time (409 otherwise).
3. A run needs about 380 generated tokens per finding left in the budget before it starts (402 otherwise).

## 1. Scan and collect on your machine
```bash
grype <image> -o json > grype.json                      # or: trivy image --format json -o trivy.json <image>
syft <image> -o cyclonedx-json > sbom.cdx.json          # optional
pip install "decosa-api @ git+<decosa-api source: on request at https://decosa.ai/contact?topic=self-host>"   # publishing soon; the collector is stdlib + httpx
python -m decosa_api.verticals.vex.collect --image <image> --findings grype.json --out evidence.json
```
The collector unpacks the image with `crane export` or `docker export` into a temp directory, searches it, and deletes it.
It fetches advisory text from OSV.dev and NVD by vulnerability id (`--offline` to use only a local store).

## 2. Triage
- `POST /vex/triage` (token). Body: `{"findings": <Grype or Trivy JSON>, "sbom"?: <CycloneDX or SPDX JSON>, "evidence"?: <bundle>,
  "image"?: {"ref", "digest", "arch"}, "select"?: {"ids"?: [...], "packages"?: [...], "severity_min"?: "high", "limit"?: 1-60},
  "author"?: "Your PSIRT", "title"?, "stream"?}` or `{"sample_id": "nginx-1.20.0"}`.
  - The default selection is known-exploited findings first, then by severity and EPSS, 25 per run (at most 60). The rest
    come back in `not_triaged`; triage them in more runs with `select.ids`.
  - Limits: 16 MB of JSON per request.
  - The JSON response has `statements` (each with `status`, `justification`, `impact_statement`, `action_statement`,
    `status_notes`, `evidence`, `checks`, `advisories`, `model` and `gate_note`), `counts`, `openvex`, `cyclonedx`,
    `envelopes` (DSSE, Ed25519), `report_md`, `record`, `record_check`, `advisories` (every record used, with its sha256),
    `not_triaged`, `receipts` and `note`.
  - With `Accept: text/event-stream` (or `"stream": true`) the events are `ready`, `advisories`, a `receipt` per model call,
    a `finding` per statement, then `result`, `budget` and `done`.
- `POST /vex/verify` (no token) `{"envelope": {...}}` checks a DSSE envelope against this server's key.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the evidence record.
- `GET /vex/info`, `GET /vex/samples` and `GET /attest/signing-key` need no token.

## Example (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"findings": json.load(open("grype.json")), "sbom": json.load(open("sbom.cdx.json")),
        "evidence": json.load(open("evidence.json")), "select": {"severity_min": "high", "limit": 40}, "author": "Example PSIRT"}
r = httpx.post(f"{API}/vex/triage", json=body, headers=H, timeout=900)
r.raise_for_status()
js = r.json()
for s in js["statements"]:
    print(f"{s['status']:>20}  {s['id']:<18} {s['package']['name']}  {s['justification'] or ''}")
json.dump(js["openvex"], open("draft.openvex.json", "w"), indent=2)      # a draft: review before publishing
json.dump(js["envelopes"], open("draft.vex.dsse.json", "w"))
json.dump(js["record"], open("vex-evidence-record.json", "w"))           # keep with the release
```
After review, `grype <image> --vex draft.openvex.json` hides the findings a not_affected or fixed statement covers
(tested with Grype 0.119; Trivy's `--vex` also reads OpenVEX but was not tested).

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa signed VEX triage: run it yourself (containers)

You are setting up Decosa's VEX triage on this machine, so images, SBOMs and scanner reports never leave it. For each
finding in a Grype or Trivy report it writes a draft VEX status with its evidence: checks run on the image (package,
version against the fix and the advisory ranges, the distro changelog, what loads the library, the advisory's programs
and settings, the platform) and one open-model judgment that code gates. `not_affected` needs positive evidence. The
output is signed OpenVEX and CycloneDX VEX drafts and a signed evidence record. Nothing is sent to Decosa's hosted API;
only vulnerability ids go to OSV.dev and NVD, and that can be switched off.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/vex-triage.zip (65 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py vex-triage` (the api image carries the same bundle under /app/rehearsal/vex-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py vex-triage --bundle vex-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "libwebp's known-exploited CVE-2023-4863 is not in the execute path (only an unloaded module loads it)", "and its status is not_affected", "OpenSSL's CVE-2022-0778 stays affected for both packages"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - set `DECOSA_VEX_OFFLINE=1` if this box must not reach the internet (then import OSV.dev's bulk export into the
     advisory store with `python -m decosa_api.verticals.vex.advisories import <zip>`);
   - keep its data on a named volume (the advisory store lives there);
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. On the host, install Grype and Syft (Apache-2.0) and `crane` (or use Docker), and get the collector: it is
   `decosa_api/verticals/vex/collect.py` in the decosa-api source and needs Python 3.11+ and `httpx`.
5. Check: `curl -fsS http://127.0.0.1:<PORT>/vex/info` lists the statuses, the justifications, the checks and the limits.
   `GET /attest/signing-key` shows this box's public key: it is what others pin to verify my VEX documents.
6. Smoke test on the bundled public sample:
   - Get a token with `POST /demo/session {"vertical":"vex-triage"}`, then send
     `POST /vex/triage {"sample_id":"nginx-1.20.0","select":{"ids":["CVE-2023-4863","CVE-2021-37600","CVE-2022-0778"]}}`.
     Expect libwebp's CVE-2023-4863 `not_affected` / `vulnerable_code_not_in_execute_path`, CVE-2021-37600 `fixed`,
     CVE-2022-0778 `affected`, and every receipt `attested`.
   - `POST /vex/verify {"envelope": <envelopes.openvex>}` should give `valid_signature: true`;
     `POST /record/verify {"record": <record>}` should give `ok: true`.
7. Then one of my own images: `grype <image> -o json > grype.json`, `syft <image> -o cyclonedx-json > sbom.json`,
   `python -m decosa_api.verticals.vex.collect --image <image> --findings grype.json --out evidence.json`, and send the
   three files to `POST /vex/triage`. Save `openvex`, `envelopes` and `record` from the answer.
8. Report back: the public key and key id, the smoke-test results, and how long the triage took.

This is a triage aid: every statement is a draft that a security engineer reviews and signs off, and it never says an
image is secure. It sees the image as shipped, not configs mounted at run time.

Off by default. Joining as a provider serves other people's requests on this GPU. Don't do it on a box that holds
internal images or SBOMs. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"vex-triage"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py vex-triage

Download the mock-data bundle (65 KB, 10 checks)expected.json

The official nginx:1.20.0 image (May 2021) scanned with Grype, and the collector's evidence bundle for three of its findings. libwebp's CVE-2023-4863 is known exploited, but only the image filter module loads libwebp and the shipped nginx.conf does not load it: it must come back not_affected with vulnerable_code_not_in_execute_path. OpenSSL's CVE-2022-0778 is in a library nginx itself links: affected. CVE-2009-4487, where NVD lists a single exact nginx version, must never be not_affected on that weak evidence. The signed OpenVEX must verify, and fail once a status is changed; the evidence record must verify.

What the rehearsal checks
  • libwebp's known-exploited CVE-2023-4863 is not in the execute path (only an unloaded module loads it)
  • and its status is not_affected
  • OpenSSL's CVE-2022-0778 stays affected for both packages
  • CVE-2009-4487 is never not_affected on an exact-version NVD entry
  • the OpenVEX document names the author
  • every statement is marked pending review
  • the signed OpenVEX verifies
  • a changed OpenVEX payload does not
  • the evidence record verifies
  • every model call has a signed receipt

Licence: Public image (Docker Official Image nginx:1.20.0). Scanner report: Grype 0.119 (Apache-2.0); evidence bundle from decosa-api's collector (Apache-2.0). Advisory text comes from the API's own store (OSV.dev and NVD, public). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa signed VEX triage: run it yourself (containers)

You are setting up Decosa's VEX triage on this machine, so images, SBOMs and scanner reports never leave it. For each
finding in a Grype or Trivy report it writes a draft VEX status with its evidence: checks run on the image (package,
version against the fix and the advisory ranges, the distro changelog, what loads the library, the advisory's programs
and settings, the platform) and one open-model judgment that code gates. `not_affected` needs positive evidence. The
output is signed OpenVEX and CycloneDX VEX drafts and a signed evidence record. Nothing is sent to Decosa's hosted API;
only vulnerability ids go to OSV.dev and NVD, and that can be switched off.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/vex-triage.zip (65 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py vex-triage` (the api image carries the same bundle under /app/rehearsal/vex-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py vex-triage --bundle vex-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "libwebp's known-exploited CVE-2023-4863 is not in the execute path (only an unloaded module loads it)", "and its status is not_affected", "OpenSSL's CVE-2022-0778 stays affected for both packages"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - set `DECOSA_VEX_OFFLINE=1` if this box must not reach the internet (then import OSV.dev's bulk export into the
     advisory store with `python -m decosa_api.verticals.vex.advisories import <zip>`);
   - keep its data on a named volume (the advisory store lives there);
   - bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check. The first start
   downloads about 20 GB of weights.
4. On the host, install Grype and Syft (Apache-2.0) and `crane` (or use Docker), and get the collector: it is
   `decosa_api/verticals/vex/collect.py` in the decosa-api source and needs Python 3.11+ and `httpx`.
5. Check: `curl -fsS http://127.0.0.1:<PORT>/vex/info` lists the statuses, the justifications, the checks and the limits.
   `GET /attest/signing-key` shows this box's public key: it is what others pin to verify my VEX documents.
6. Smoke test on the bundled public sample:
   - Get a token with `POST /demo/session {"vertical":"vex-triage"}`, then send
     `POST /vex/triage {"sample_id":"nginx-1.20.0","select":{"ids":["CVE-2023-4863","CVE-2021-37600","CVE-2022-0778"]}}`.
     Expect libwebp's CVE-2023-4863 `not_affected` / `vulnerable_code_not_in_execute_path`, CVE-2021-37600 `fixed`,
     CVE-2022-0778 `affected`, and every receipt `attested`.
   - `POST /vex/verify {"envelope": <envelopes.openvex>}` should give `valid_signature: true`;
     `POST /record/verify {"record": <record>}` should give `ok: true`.
7. Then one of my own images: `grype <image> -o json > grype.json`, `syft <image> -o cyclonedx-json > sbom.json`,
   `python -m decosa_api.verticals.vex.collect --image <image> --findings grype.json --out evidence.json`, and send the
   three files to `POST /vex/triage`. Save `openvex`, `envelopes` and `record` from the answer.
8. Report back: the public key and key id, the smoke-test results, and how long the triage took.

This is a triage aid: every statement is a draft that a security engineer reviews and signs off, and it never says an
image is secure. It sees the image as shipped, not configs mounted at run time.

Off by default. Joining as a provider serves other people's requests on this GPU. Don't do it on a box that holds
internal images or SBOMs. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsSigned VEX triage on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • One typed judgment per finding the rules do n...: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • The collector: decosa VEX checks and collector (code, no model). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Signed VEX triage, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Signed VEX triage on my hardware

Fetch https://decosa.ai/prompts/vex-triage-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=vex-triage)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- One typed judgment per finding the rules do n...: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- The collector: decosa VEX checks and collector (code, no model), CPU

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/vex-triage-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 14 s · ~$0.014 per run · 12 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing, DECOSA_VEX_OFFLINE=1; torn down after

Measured cost to run: about $0.12 per 100 findings (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The assembly prompt's smoke test passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): libwebp CVE-2023-4863 not_affected (not in execute path), CVE-2022-0778 affected, CVE-2009-4487 under investigation, all receipts attested, envelope and record verified, 2.0 s; the rehearsal bundle passed 10/10. The collector's --image path (crane) found that absolute symlinks were dropped when unpacking; fixed in b5a465e and re-checked (5,370 files). Model-server startup was not re-run.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured against one vendor's VEX (Canonical) on two Ubuntu releases, with constructed probe findings; not against Red Hat or Debian data or a security engineer's review.
  • Recall of not_affected is about one in three: distro VEX often rests on maintainer judgment this tool does not make.
  • Language packages (PyPI, npm, Go) use the same OSV range check but were not measured against labels; there is no language-level reachability.
  • NVD allows 5 requests per 30 s without a key, so the hosted service fetches few new NVD records per run; findings then rest on OSV and the scanner's own data.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the signed vex triage API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the collector, the checks, rules mode and the signing run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

A Grype or Trivy report in; a VEX status per CVE with the evidence from the image, signed OpenVEX and CycloneDX VEX out.

For product security teams and software vendors. Give it a scanner report for a container image, its SBOM and an evidence bundle from a small collector that runs next to the image. For each finding, code checks the package, the version against the fix and the OSV and NVD ranges, the distro changelog, what actually loads the library, the settings and programs the advisory names, and the platform; an open model writes a VEX status with quoted evidence, and code gates it. Not affected needs a strong check; everything else stays affected, fixed or under investigation. Drafts for a security engineer to sign off.

Deployment
Hosted or self-host
Regulatory
Checked 26 Sep 2026. Statuses and justifications follow CISA's Minimum Requirements for Vulnerability Exploitability eXchange (VEX) (21 Apr 2023, https://www.cisa.gov/resources-tools/resources/minimum-requirements-vulnerability-exploitability-exchange-vex) and OpenVEX v0.2.0 (https://github.com/openvex/spec); CycloneDX 1.6 VEX uses its own justification vocabulary (mapping in GET /vex/info). EU Cyber Resilience Act, Regulation (EU) 2024/2847: vulnerability and incident reporting (Art. 14) applies from 11 Sep 2026 and the main obligations, including documenting components with a software bill of materials (Annex I Part II), from 11 Dec 2027 (Art. 71(2)); we could not re-read the Official Journal text on 26 Sep 2026, so treat these dates as from our 26 Sep incident-pack check, not independently verified here. VEX is a way to state exploitability; this tool drafts statements and never says an image is secure, and it makes no CRA or other compliance determination. Model licence: Apache-2.0 (Qwen3.8-27B). Advisory sources: OSV.dev (per-record licence; GitHub advisories CC BY 4.0) and NVD (US government work).
Architecture
Text description

On your machine, the image is scanned with Grype or Trivy and Syft, and the collector searches the unpacked image and writes an evidence bundle; the image never leaves. decosa-api reads the report, the SBOM and the bundle, looks up advisory text in a versioned store filled from OSV.dev and NVD, and runs checks in code. Findings the rules settle are marked fixed without a model call; for the rest Qwen3.8-27B (Apache-2.0) picks a status with quoted evidence, and code gates it so that not affected stands only on a strong check. Outputs: signed OpenVEX and CycloneDX VEX (DSSE, Ed25519) and a signed evidence record, reviewed by a security engineer. Hosted calls get gateway-signed receipts; self-hosted, the box signs its own.

Architecture

At a glance

What it gives you
A VEX status per finding with its evidence (the checks and the advisory lines it rests on), signed OpenVEX and CycloneDX VEX drafts in DSSE envelopes, a Markdown report and a signed evidence record listing every advisory's hash.
What it does not do
It does not scan (it triages a Grype or Trivy report), trace calls inside programs, see configs mounted at run time or dlopen by path, or decide attacker control or mitigations. It never says an image is secure; a security engineer signs off.
Data retention
Nothing kept but advisory text: reports, SBOMs and evidence bundles live in memory for the request; logs carry counts only. The image never reaches the API: the collector runs on your machine.
What leaves the box
Hosted: the scanner report, SBOM and evidence bundle (search hits, loader chains, config excerpts with secret-looking lines redacted) go to decosa-api, and the prompts to Qwen3.8-27B through our gateway, whose receipts hold hashes. Both routes: vulnerability ids to OSV.dev and NVD, which you can switch off.
Round trip
Tested: grype --vex with the signed OpenVEX from the nginx demo hides exactly the four findings it proves not affected and keeps the other 546.
Typical run cost
One model call per finding the rules do not settle: about a cent for the nginx sample at the gateway list price. Rules mode costs no model tokens.
Findings per run
25 by default, 60 hosted (known-exploited first, then severity and EPSS); self-host raises it with DECOSA_VEX_TRIAGE_MAX.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    rules only, no GPU

    The same checks and gates with no model call (mode "rules"): the same statuses on the held-out set, but no written impact statement beyond the check's own text, and nothing routed to under investigation.

    Models
    • decosa VEX checks and collector (code, no model)
    Hardware
    Any CPU
    Quality evidence
    • not_affected precision against Canonical's VEX, held-out jammy set159/164decosa-api docs/evals/vex-triage.md, 2026-09-26 (rules-only baseline on the same run)
    • false not_affected on vendor-affected findings, held-out5/378decosa-api docs/evals/vex-triage.md, 2026-09-26
    • not_affected recall, held-out159/474decosa-api docs/evals/vex-triage.md, 2026-09-26
    Latency
    measured: the checks for hundreds of findings take seconds on CPU
    Verification
    No proof yetSelf-host onlyNo model call, so no receipts; the VEX documents and the record are signed by the server's key.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    The checks plus one Qwen3.8-27B judgment per finding the rules do not settle: a written impact statement with quotes, and unclear cases sent to a person. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    • decosa VEX checks and collector (code, no model)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • not_affected precision against Canonical's VEX, held-out jammy set (912 findings, run once)154/159decosa-api docs/evals/vex-triage.md, measured on our server 2026-09-26, gateway route; code frozen on the focal dev set
    • false not_affected on vendor-affected findings (the risky error), held-out5/378; all five checks hold on the image (other platform, or a program from another package) but Canonical scores the source packagedecosa-api docs/evals/vex-triage.md, 2026-09-26
    • not_affected recall / under_investigation rate, held-out154/474 / 106/852decosa-api docs/evals/vex-triage.md, 2026-09-26
    • supporting checks re-done with readelf, find and grep159/161 hold, 0 fail (2 could not be parsed)decosa-api docs/evals/vex-triage.md, scripts/vex_evidence_audit.py, 2026-09-26
    • labels from Red Hat, Debian or a security engineernot measured yet
    Latency
    measured on our server under a shared gateway: seconds for the nginx sample; many minutes for hundreds of findings with calls in parallel
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
One typed judgment per finding the rules do not settle: a VEX status with quoted evidence from the advisory lines and the checks, and the impact statement. Code gates every answer.Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
The collector (on your machine) and the checks in decosa-api: package in the SBOM, version against the fix and the OSV/NVD ranges, the Debian changelog, ELF loader chains, the advisory's programs and settings, the platform; the evidence gates; OpenVEX, CycloneDX VEX and DSSE signing.decosa VEX checks and collector (code, no model)
0 GBNo proof yetIn the hosted demo

Measured

Does it only say not affected when it can show why?

Grype findings on two public Ubuntu images, scored against Canonical's own OpenVEX for the same CVE and package. Canonical's not_affected statements for packages in the image were added as probes, shaped like an upstream-matching scanner's findings. Everything was tuned on ubuntu:focal-20210416; ubuntu:jammy-20220421 was run once with the code frozen.

not_affected that Canonical agrees with
154 of 159held-out jammy set; dev 19 of 24
Vendor-affected findings wrongly marked not_affected
5 of 378all five checks hold on the image: 32-bit, PowerPC or Windows only, or a program from another package
Canonical's not_affected found
154 of 474the rest rest on a maintainer's judgment this tool never makes
nginx:1.20.0 findings closed with strong evidence
168 of 550rules mode; mostly libraries only unloaded modules pull in; no vendor labels for this image

Where it disagrees

Its not_affected answers are about the image, and a distro's VEX is about the source package: an AES-OCB bug that only affects 32-bit x86 is not in the execute path of an amd64 image, but Canonical still ships the fix and calls the package affected. All five disagreements on the held-out set are of that kind.

What the model adds

Little to the decisions: the rules alone reach the same precision (159 of 164). The model writes the impact statement with quotes and sends unclear cases to under investigation; the gates turned 101 of its unsupported not_affected proposals into under investigation.

Source: decosa-api docs/evals/vex-triage.md, 2026-09-26

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    The checks, the gates, the advisory store, DSSE signing and the HTTP API (/vex/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network; not needed in rules mode.

  • collector (CLI, on your machine)

    python -m decosa_api.verticals.vex.collect: unpacks the image, searches it, writes the evidence bundle. The image never leaves the machine.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted Qwen3.8-27B runs on one of these cards on our server.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context.

  • CPU only Fits

    Rules mode (no model), the collector, signing and verification need no GPU; measured on the held-out set with the same precision as the model.

Latency per lane

  • nginx sample, 12 findings, busy shared gateway14.3 s

    Measuredmeasured on our server 2026-09-26, pre-release server, gateway route (12 model calls)

  • 912 findings, 8 calls in parallel1372.0 s

    Measureddecosa-api docs/evals/vex-triage.md, held-out jammy run, 2026-09-26

  • collector on nginx:1.20.0 (550 findings, 5,370 files)1.9 s

    Measuredmeasured on our server 2026-09-26, rootfs already unpacked

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

vex-triage/assemble-prompt.md188 lines
# Assemble Decosa signed VEX triage on this machine

You are setting up self-hosted VEX triage on this Linux machine for a product security team. It reads a scanner report
(Grype or Trivy JSON) for a container image, optionally its SBOM, and an evidence bundle that a collector writes by
searching the unpacked image, and returns:
- one draft VEX statement per finding: `not_affected` with a CISA justification, `affected` with an action, `fixed`, or
  `under_investigation`, each with the checks and quotes behind it;
- signed OpenVEX and CycloneDX VEX documents (DSSE envelopes, Ed25519); Grype reads the OpenVEX back with `--vex` (tested);
- a signed, hash-chained evidence record of the inputs, every advisory used and every statement.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Images, SBOMs and scanner reports stay on this machine. The only outbound calls are advisory lookups by vulnerability id
  (OSV.dev and NVD), and those can be switched off for an air-gapped box.
- This is a triage aid. Every statement is a draft that a security engineer reviews and signs off; it never says an
  image is secure.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/vex-triage.zip (65 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py vex-triage` (the api image carries the same bundle under /app/rehearsal/vex-triage/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py vex-triage --bundle vex-triage.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "libwebp's known-exploited CVE-2023-4863 is not in the execute path (only an unloaded module loads it)", "and its status is not_affected", "OpenSSL's CVE-2022-0778 stays affected for both packages"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

On the host (not in a container): Grype or Trivy, Syft, `crane` (or Docker) and Python 3.11+ for the collector.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`; on 48 GB also
     `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk, plus room to unpack the images you triage.
4. Install the scanner tools if missing (all Apache-2.0): Grype and Syft from `github.com/anchore/{grype,syft}` releases,
   `crane` from `github.com/google/go-containerregistry` releases (or Trivy from `github.com/aquasecurity/trivy`).

## 2. Get the images and the collector

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
If a pull fails, build from source once `decosa-api` is published: clone it, run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from
its compose file. If neither works, stop and tell me.

The collector is part of the same source. In the clone: `python3 -m venv .venv && .venv/bin/pip install httpx`. It
needs nothing else (`python -m decosa_api.verticals.vex.collect --help`).

## 3. Write the compose file

Create `~/decosa-vex/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these VEX documents, e.g. Example Corp PSIRT>"
```

Create `~/decosa-vex/docker-compose.yml` with exactly these services:

```yaml
name: decosa-vex
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "400000"        # per session; about 380 generated tokens per finding
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_VEX_TRIAGE_MAX: "200"              # findings per request (the hosted service allows 60)
      DECOSA_VEX_OFFLINE: "0"                   # "1" for air-gapped: advisories only from the local store
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` exactly as written: a root-owned host bind mount makes the API fail on
`/data/keys.sqlite`. The advisory store lives at `/data/vex/advisories` in that volume.

Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy (the LLM takes 5-10 minutes
the first time). `curl -s localhost:8445/healthz` should show `"llm": true`.

## 4. Smoke test on the bundled sample

The API ships a sample: the public `nginx:1.20.0` image's Grype report, SBOM and evidence bundle, with its advisories.

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"vex-triage"}' | jq -r .token)
curl -s $API/vex/triage -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"nginx-1.20.0","select":{"ids":["CVE-2023-4863","CVE-2021-37600","CVE-2022-0778"]}}' > /tmp/vex.json
jq -r '.statements[] | "\(.status)\t\(.justification // "")\t\(.id)\t\(.package.name)"' /tmp/vex.json
jq '{envelope: .envelopes.openvex}' /tmp/vex.json | curl -s $API/vex/verify -H 'content-type: application/json' -d @- | jq '{valid_signature, statements}'
jq '{record}' /tmp/vex.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if CVE-2023-4863 (libwebp6) is `not_affected` / `vulnerable_code_not_in_execute_path`, CVE-2021-37600 is `fixed`,
CVE-2022-0778 is `affected`, every receipt is `attested`, the envelope verifies and the record verifies.

## 5. Triage your own image

```bash
IMG=registry.example.com/app:1.4.2
grype $IMG -o json > grype.json
syft $IMG -o cyclonedx-json > sbom.cdx.json
.venv/bin/python -m decosa_api.verticals.vex.collect --image $IMG --findings grype.json --out evidence.json
jq -n --slurpfile f grype.json --slurpfile s sbom.cdx.json --slurpfile e evidence.json \
  '{findings: $f[0], sbom: $s[0], evidence: $e[0], select: {severity_min: "high", limit: 100}, author: "Example Corp PSIRT"}' \
  | curl -s $API/vex/triage -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- > vex-run.json
jq '.openvex' vex-run.json > draft.openvex.json      # review, then:  grype $IMG --vex draft.openvex.json
```

Air-gapped: import advisories first (`python -m decosa_api.verticals.vex.advisories import osv-dump.zip`, from OSV.dev's
bulk export), run the collector with `--offline`, and set `DECOSA_VEX_OFFLINE=1`.

## 6. Point tools at the local API

- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /vex/triage` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /vex/info` lists the
  statuses, justifications, checks and limits.
- Keep the signed OpenVEX, the envelopes and the evidence record with the release. Anyone can re-check them with
  `POST /vex/verify` and `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts (advisory text, package names and your evidence) to the hosted Decosa
API. Leave
it off for internal images.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

2 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A Grype or Trivy report in; a VEX status per CVE with the evidence from the image, signed OpenVEX and CycloneDX VEX out.
Who it's for
Product security (PSIRT) and AppSec teams at software vendors shipping container images, and platform teams that own base images.
Where it runs
Hosted or self-host (air-gapped with a local advisory store)
Key numbers
  • 154 / 159 not_affected precision against Canonical's VEX (test split, n = 159)
  • 5 / 378 False not_affected on vendor-affected findings (test split, n = 378)
  • 154 / 474 not_affected recall (test split, n = 474)
  • 14.3 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Hosted or self-host (air-gapped with a local advisory store)
Checks
Receipt per model call; every quote found in the advisory or check it cites; not_affected only on a strong code check; DSSE-signed VEX documents; signed hash-chained evidence record
Output
Signed record or verdict · Structured data
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

When does it say a CVE is not affected?

Only when a check of the image shows why: the version is outside every affected range, the program the advisory names is not in the image, no program loads the library (or only a plugin module no config enables), the setting the advisory needs is off, or the advisory limits the bug to another platform. On a held-out Ubuntu set, Canonical agreed with 154 of its 159 not affected answers.

Does my container image leave my machine?

No. A small collector runs next to the image, unpacks it, searches it and writes an evidence bundle: file counts, search hits, library loader chains and config excerpts with secret-looking lines redacted. You send that, the scanner report and optionally the SBOM. Self-hosted, only vulnerability ids go to OSV.dev and NVD, and that can be switched off.

Can Grype use the VEX it writes?

Yes, tested: feeding the signed OpenVEX from the nginx:1.20.0 demo back to grype --vex hid exactly the four findings it proved not affected and kept the other 546. Trivy reads OpenVEX too, but we did not test that here.

Will it tell me my image is secure?

No. Every statement is a draft for a security engineer to review and sign off. It sees the image as shipped, not configs mounted at run time, and it does not trace calls inside programs.

What does the model do?

Qwen3.8-27B writes each status with quotes from the advisory and the checks, and code gates it: a not affected without a strong check becomes under investigation. On the held-out set the rules alone reached the same precision, so a rules-only mode with no GPU ships too.

How is it different from NVIDIA's Vulnerability Analysis blueprint?

The blueprint is a kit that defaults to a 550B model through NVIDIA's API and web search, with no signature. This runs on one card or on CPU, borrows the blueprint's Apache-2.0 version and changelog checks, and signs the VEX and an evidence record.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Signed VEX triage

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.