Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Signed VEX triage (63): eval

Run on our server on 26 Sep 2026, gateway route (Qwen3.8-27B through the model gateway, every call receipted), shared gateway under load. Runner: scripts/vex_eval.py; independent evidence re-check: scripts/vex_evidence_audit.py. Working data (images, scans, Canonical's feed, the advisory store) lives in <internal path> on our server and is not shipped.

What was measured, on what

Images. Two public Ubuntu images, old enough to carry many known CVEs:

  • ubuntu:focal-20210416 is the dev set. Every prompt and code change was made while looking at it.
  • ubuntu:jammy-20220421 is the held-out test set. The code was frozen at commit 38df889, then this set was run once. Nothing was changed after seeing its results.

Scanner findings. Each image was scanned with Grype 0.119, whose database is dated 26 Sep 2026, and given a Syft SBOM. The collector unpacked each image with crane export and built the evidence bundle.

Labels. The labels are Canonical's own OpenVEX feed (vex-all.tar.xz, security-metadata.canonical.com, downloaded 26 Sep 2026). A finding is labelled from the statement for the same CVE, binary package and release (including esm-infra):

  • vendor fixed at a version newer than the image's is labelled affected, because the image still has the vulnerable version;
  • vendor not_affected is labelled not_affected, with its justification;
  • vendor under_investigation is reported but not scored.

Probes. Grype already applies Canonical's data, so its findings are almost all vendor-affected. They test the risky direction: a false not_affected. To test the other direction, every Canonical not_affected statement for a package in the image was added as a probe. A probe is shaped like a Grype match with no distro fix data, which is how a CPE- or upstream-matching scanner would report it. Nothing from Canonical's statement reaches the triage.

Leakage. The triage reads only these, and never Ubuntu's OSV records (UBUNTU-CVE-*) or USNs:

  • the OSV.dev CVE record and its GHSA aliases;
  • the NVD record;
  • the scanner's own text.

Baseline. "Rules only" means the same checks with no model. This is also the API's mode: "rules".

Numbers

dev: focal (tuned on) test: jammy (held out, run once)
Findings triaged (scanner + probes) 458 (408 + 50) 912 (438 + 474)
Scored (vendor affected / vendor not_affected) 321 (271 / 50) 852 (378 / 474)
not_affected precision 19 / 24 154 / 159 (96.9%)
False not_affected on vendor-affected findings (the risky error) 5 / 271 5 / 378 (1.3%)
False fixed on vendor-affected findings 0 / 271 0 / 378
not_affected recall (vendor not_affected found) 19 / 50 154 / 474 (32.5%)
Justification equal to Canonical's, when both say not_affected 19 / 19 151 / 154
affected recall 252 / 271 365 / 378
under_investigation rate 25 / 321 (7.8%) 106 / 852 (12.4%)
Status agreement, all scored 271 / 321 519 / 852
Vendor under_investigation (not scored): ours 119 affected, 4 under inv. 54 affected, 4 under inv., 2 not_affected

Rules only, same checks and no model, test set:

  • not_affected precision 159 / 164;
  • false not_affected 5 / 378;
  • recall 159 / 474;
  • agreement 532 / 852;
  • under_investigation 0.

The rules make the same five risky errors as the model and agree with Canonical slightly more often.

What the model's under_investigation answers cover on the test set: 98 are vendor-not_affected, 8 vendor-affected and 4 vendor-under-investigation. The rules would have said affected for almost all of them.

Evidence validity

  • Independent re-check. For every not_affected statement, each supporting check was re-done on the unpacked image with other tools:
    • readelf -d rebuilt the library-loader graph;
    • find -name searched for the named programs;
    • grep searched the configs;
    • file confirmed the image's architecture;
    • packaging.version recomputed the NVD bounds.
  • Test set: 159 / 161 checks hold, 0 fail. The other 2 could not be parsed by the second implementation: NVD versions r118 and 1.19.7ubuntu3. Dev: 24 / 26 hold, 0 fail, 2 unparsed. Output: eval/<image>.<set>.audit.json.
  • Model quotes. On the test set, 2,386 quotes were found where the model said and 204 (7.9%) were not. Those were dropped and do not count as evidence. Dev: 1,246 / 114, after a fix that strips the [strong] label the model copied into quotes.
  • Read by hand. The building agent read 8 random not_affected and 3 affected statements from dev. All impact statements matched their checks. Examples:
    • "Upstream 8.30 is outside every NVD range for coreutils (versionEndIncluding 8.29)";
    • "limits the issue to 32-bit platforms; the image is amd64".

The five risky errors (test), read one by one

All five supporting checks hold on the image. They disagree with Canonical because Canonical writes about the source package, while this tool writes about the image:

  • CVE-2022-2097 (libssl3): AES OCB "for 32-bit x86 platforms". The image is amd64. Canonical: fixed later.
  • CVE-2023-6129 (libssl3): POLY1305 "on PowerPC CPU based platforms". The image is amd64.
  • CVE-2026-8376 (perl-base): "on 32-bit builds".
  • CVE-2023-47039 (perl-base): Perl on Windows.
  • CVE-2026-42250 (libbz2-1.0): the bug is in bzip2recover, which is in the bzip2 package, not in the image.

They are still counted as errors above. A reviewer should know that "other platform" answers are this tool's most frequent disagreement with a distro's VEX.

Cost and speed

  • Test run: 912 findings in 1,372 s with 8 parallel calls on a busy shared gateway.
  • Smoke (4 findings): 13,233 tokens, $0.0048 at the gateway list price. That is about 3,300 tokens and $0.0012 per model call, one call per finding that the rules do not settle.
  • Hosted sample runs on the pre-release server: nginx, 12 findings, 14.3 s; python, 12 findings, 13.8 s.

Whole reports (no labels): how much noise the checks remove

In rules mode, with the same checks the model sees:

  • nginx:1.20.0: 168 of 550 Grype findings get a not_affected with strong evidence. 164 are vulnerable_code_not_in_execute_path, mostly libtiff, libwebp, libexpat, libxml2, libgd, libX11 and freetype, which only the image-filter and XSLT dynamic modules load, and the shipped nginx.conf loads neither. 4 are component_not_present.
  • python:3.10.0-slim-bullseye: 10 of 637.

No vendor labels exist for these images (Debian publishes no OpenVEX), so these counts are not accuracy.

Round trip into the scanner

The OpenVEX from the recorded nginx-1.20.0 run was fed back to Grype with grype registry:nginx:1.20.0 --vex doc.openvex.json. Grype moved exactly the 4 not_affected findings (libwebp, libexpat and two libtiff) to ignoredMatches and kept the other 546.

This needed one fix, found while testing. The image identity must be the repository digest, not Grype's manifest digest, and the pkg:oci product must have no tag qualifier. Trivy's --vex was not tested.

Expected properties of the sample runs (for the rehearsal kit)

  1. nginx-1.20.0, CVE-2023-4863 (libwebp6, CISA KEV): not_affected / vulnerable_code_not_in_execute_path. The loader chain is libwebp.so.6 <- libgd.so.3 <- ngx_http_image_filter_module.so, and no config file names that module.
  2. nginx-1.20.0-image-filter (the same image with a load_module line added): CVE-2023-4863 becomes affected.
  3. CVE-2022-0778 (libssl1.1 and openssl): affected in both nginx samples, because nginx links libssl.
  4. CVE-2009-4487 (nginx): never not_affected. NVD lists only 0.7.64, so the check is weak and the answer is under_investigation.
  5. python-3.10.0-slim, CVE-2022-2097 (32-bit x86 only): not_affected / vulnerable_code_not_in_execute_path. CVE-2023-4911 (glibc): under_investigation, because the only evidence is weak.
  6. Every model call has a signed receipt. The OpenVEX DSSE envelope verifies at /vex/verify and fails after any payload change. The evidence record verifies at /record/verify.

rehearsal/vex-triage/ checks 1, 3, 4 and 6 (10 checks, passed 10/10 on the pre-release server).

Honest limits

  • Labels. They come from one vendor (Canonical) and two releases. Canonical's statements are package-level and include judgment calls, such as "I can't see how this could be exploited" and "disputed". This tool never makes those calls, which explains most of the low not_affected recall (32.5%). No Red Hat, Debian or Chainguard labels were used. Chainguard's advisories are CC BY-NC-ND, so they were left out of a commercial eval.
  • Probes are constructed findings. A real CPE-matching scanner would flag some of them and not others, so the precision numbers depend on the mix. Read the per-class counts, not the ratio alone.
  • The model adds little to the decisions. Rules only has the same precision and slightly higher agreement. The model writes the impact statements and sends unclear cases to a person (under_investigation), and the gates stop it from over-claiming: on the test set, 101 not_affected proposals without strong evidence became under_investigation. A buyer could run mode: "rules" (no GPU) and get about the same statuses.
  • Checks see the image as shipped. Configs mounted at run time and dlopen by path are not seen. Library use is read from ELF links only. There is no call-graph reachability inside programs, and no language-level reachability (Python imports, Go symbols).
  • Calibration. It is fit on dev and separates little: the model says 95 on 308 of 320. Treat probability as informational.
  • Distro packages only. Only distro packages are evaluated. Language packages (PyPI, npm, Go) use OSV ranges in the same code path but were not measured against labels.
  • Who judged. The building agent chose the probe method and read the statements. There was no independent security engineer.