Skip to content
decosa

63 · Software and AI ops · Compliance and trust · live

Signed VEX triage

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 26 Sep 2026Eval write-up (decosa-api, access required)

  • not_affected precision against Canonical's VEX154 / 159test splitn = 159held-out ubuntu:jammy-20220421, run once with the code frozen
  • False not_affected on vendor-affected findings5 / 378test splitn = 378all five supporting checks hold on the image (other platform, or a program from another package)
  • not_affected recall154 / 474test splitn = 474Canonical's not_affected statements for packages in the image, added as probes
  • under_investigation rate106 / 852test splitn = 85298 of the 106 are findings Canonical calls not_affected
  • Justification equal to Canonical's151 / 154test splitn = 154when both say not_affected
  • Supporting checks that hold when re-done with other tools159 / 161test splitn = 161readelf, find, grep, file, packaging.version; 0 fail, 2 could not be parsed
  • Rules only (no model): not_affected precision159 / 164test splitn = 164same checks, no model call
  • not_affected precision, dev set19 / 24dev (tuned on)n = 24ubuntu:focal-20210416, where prompts and code were tuned

Dataset

Grype 0.119 findings on ubuntu:focal-20210416 (dev, 458 findings) and ubuntu:jammy-20220421 (test, 912), plus Canonical's not_affected statements for packages in each image as probe findings; labels from Canonical's OpenVEX feed, 26 Sep 2026.

Caveats

  • One vendor's labels (Canonical) on two releases; Canonical writes about source packages and this tool about the image, which explains all five risky errors on the test set.
  • Probe findings are constructed, so precision depends on the mix; read the per-class counts.
  • The model adds little to the decisions: rules only reach the same precision. The building agent read the statements; no independent security engineer.
  • Distro packages only; language packages were not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
14 s
Receipts
12
Model calls
n/a
Tokens
n/a
Cost per run
$0.014

Self-host verification

Verified on 26 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing, DECOSA_VEX_OFFLINE=1; torn down after

The assembly prompt's smoke test passed against the already-running local Qwen3.8-27B vLLM (network_mode host instead of the compose llm service): libwebp CVE-2023-4863 not_affected (not in execute path), CVE-2022-0778 affected, CVE-2009-4487 under investigation, all receipts attested, envelope and record verified, 2.0 s; the rehearsal bundle passed 10/10. The collector's --image path (crane) found that absolute symlinks were dropped when unpacking; fixed in b5a465e and re-checked (5,370 files). Model-server startup was not re-run.

Rehearsal bundle: vex-triage.zip (65 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
  • Measured against one vendor's VEX (Canonical) on two Ubuntu releases, with constructed probe findings; not against Red Hat or Debian data or a security engineer's review.
  • Recall of not_affected is about one in three: distro VEX often rests on maintainer judgment this tool does not make.
  • Language packages (PyPI, npm, Go) use the same OSV range check but were not measured against labels; there is no language-level reachability.
  • NVD allows 5 requests per 30 s without a key, so the hosted service fetches few new NVD records per run; findings then rest on OSV and the scanner's own data.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • One typed judgment per finding the rules do not settle: a VEX status with quoted evidence from the advisory lines and the checks, and the impact statement. Code gates every answer.Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • The collector (on your machine) and the checks in decosa-api: package in the SBOM, version against the fix and the OSV/NVD ranges, the Debian changelog, ELF loader chains, the advisory's programs and settings, the platform; the evidence gates; OpenVEX, CycloneDX VEX and DSSE signing.decosa VEX checks and collector (code, no model)AGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Does it only say not affected when it can show why?

  • not_affected that Canonical agrees with: 154 of 159 (held-out jammy set; dev 19 of 24)
  • Vendor-affected findings wrongly marked not_affected: 5 of 378 (all five checks hold on the image: 32-bit, PowerPC or Windows only, or a program from another package)
  • Canonical's not_affected found: 154 of 474 (the rest rest on a maintainer's judgment this tool never makes)
  • nginx:1.20.0 findings closed with strong evidence: 168 of 550 (rules mode; mostly libraries only unloaded modules pull in; no vendor labels for this image)

Source: decosa-api docs/evals/vex-triage.md, 2026-09-26

Lite · rules only, no GPU (3)
  • not_affected precision against Canonical's VEX, held-out jammy set: 159/164decosa-api docs/evals/vex-triage.md, 2026-09-26 (rules-only baseline on the same run)
  • false not_affected on vendor-affected findings, held-out: 5/378decosa-api docs/evals/vex-triage.md, 2026-09-26
  • not_affected recall, held-out: 159/474decosa-api docs/evals/vex-triage.md, 2026-09-26
Standard · the hosted demo, one 96 GB card (5)
  • not_affected precision against Canonical's VEX, held-out jammy set (912 findings, run once): 154/159decosa-api docs/evals/vex-triage.md, measured on our server 2026-09-26, gateway route; code frozen on the focal dev set
  • false not_affected on vendor-affected findings (the risky error), held-out: 5/378; all five checks hold on the image (other platform, or a program from another package) but Canonical scores the source packagedecosa-api docs/evals/vex-triage.md, 2026-09-26
  • not_affected recall / under_investigation rate, held-out: 154/474 / 106/852decosa-api docs/evals/vex-triage.md, 2026-09-26
  • supporting checks re-done with readelf, find and grep: 159/161 hold, 0 fail (2 could not be parsed)decosa-api docs/evals/vex-triage.md, scripts/vex_evidence_audit.py, 2026-09-26
  • labels from Red Hat, Debian or a security engineer: not measured yet

How we measure · All tools