Signed VEX triage (63): eval
Run on our server on 26 Sep 2026, gateway route (Qwen3.8-27B through the model gateway, every call receipted), shared gateway
under load. Runner: scripts/vex_eval.py; independent evidence re-check: scripts/vex_evidence_audit.py. Working data
(images, scans, Canonical's feed, the advisory store) lives in <internal path> on our server and is not
shipped.
What was measured, on what
Images. Two public Ubuntu images, old enough to carry many known CVEs:
ubuntu:focal-20210416is the dev set. Every prompt and code change was made while looking at it.ubuntu:jammy-20220421is the held-out test set. The code was frozen at commit38df889, then this set was run once. Nothing was changed after seeing its results.
Scanner findings. Each image was scanned with Grype 0.119, whose database is dated 26 Sep 2026, and given a Syft
SBOM. The collector unpacked each image with crane export and built the evidence bundle.
Labels. The labels are Canonical's own OpenVEX feed (vex-all.tar.xz, security-metadata.canonical.com, downloaded
26 Sep 2026). A finding is labelled from the statement for the same CVE, binary package and release (including
esm-infra):
- vendor
fixedat a version newer than the image's is labelled affected, because the image still has the vulnerable version; - vendor
not_affectedis labelled not_affected, with its justification; - vendor
under_investigationis reported but not scored.
Probes. Grype already applies Canonical's data, so its findings are almost all vendor-affected. They test the risky
direction: a false not_affected. To test the other direction, every Canonical not_affected statement for a package in
the image was added as a probe. A probe is shaped like a Grype match with no distro fix data, which is how a CPE- or
upstream-matching scanner would report it. Nothing from Canonical's statement reaches the triage.
Leakage. The triage reads only these, and never Ubuntu's OSV records (UBUNTU-CVE-*) or USNs:
- the OSV.dev CVE record and its GHSA aliases;
- the NVD record;
- the scanner's own text.
Baseline. "Rules only" means the same checks with no model. This is also the API's mode: "rules".
Numbers
| dev: focal (tuned on) | test: jammy (held out, run once) | |
|---|---|---|
| Findings triaged (scanner + probes) | 458 (408 + 50) | 912 (438 + 474) |
| Scored (vendor affected / vendor not_affected) | 321 (271 / 50) | 852 (378 / 474) |
| not_affected precision | 19 / 24 | 154 / 159 (96.9%) |
| False not_affected on vendor-affected findings (the risky error) | 5 / 271 | 5 / 378 (1.3%) |
| False fixed on vendor-affected findings | 0 / 271 | 0 / 378 |
| not_affected recall (vendor not_affected found) | 19 / 50 | 154 / 474 (32.5%) |
| Justification equal to Canonical's, when both say not_affected | 19 / 19 | 151 / 154 |
| affected recall | 252 / 271 | 365 / 378 |
| under_investigation rate | 25 / 321 (7.8%) | 106 / 852 (12.4%) |
| Status agreement, all scored | 271 / 321 | 519 / 852 |
| Vendor under_investigation (not scored): ours | 119 affected, 4 under inv. | 54 affected, 4 under inv., 2 not_affected |
Rules only, same checks and no model, test set:
- not_affected precision 159 / 164;
- false not_affected 5 / 378;
- recall 159 / 474;
- agreement 532 / 852;
- under_investigation 0.
The rules make the same five risky errors as the model and agree with Canonical slightly more often.
What the model's under_investigation answers cover on the test set: 98 are vendor-not_affected, 8 vendor-affected and 4
vendor-under-investigation. The rules would have said affected for almost all of them.
Evidence validity
- Independent re-check. For every not_affected statement, each supporting check was re-done on the unpacked image
with other tools:
readelf -drebuilt the library-loader graph;find -namesearched for the named programs;grepsearched the configs;fileconfirmed the image's architecture;packaging.versionrecomputed the NVD bounds.
- Test set: 159 / 161 checks hold, 0 fail. The other 2 could not be parsed by the second implementation: NVD
versions
r118and1.19.7ubuntu3. Dev: 24 / 26 hold, 0 fail, 2 unparsed. Output:eval/<image>.<set>.audit.json. - Model quotes. On the test set, 2,386 quotes were found where the model said and 204 (7.9%) were not. Those were
dropped and do not count as evidence. Dev: 1,246 / 114, after a fix that strips the
[strong]label the model copied into quotes. - Read by hand. The building agent read 8 random not_affected and 3 affected statements from dev. All impact
statements matched their checks. Examples:
- "Upstream 8.30 is outside every NVD range for coreutils (versionEndIncluding 8.29)";
- "limits the issue to 32-bit platforms; the image is amd64".
The five risky errors (test), read one by one
All five supporting checks hold on the image. They disagree with Canonical because Canonical writes about the source package, while this tool writes about the image:
- CVE-2022-2097 (libssl3): AES OCB "for 32-bit x86 platforms". The image is amd64. Canonical: fixed later.
- CVE-2023-6129 (libssl3): POLY1305 "on PowerPC CPU based platforms". The image is amd64.
- CVE-2026-8376 (perl-base): "on 32-bit builds".
- CVE-2023-47039 (perl-base): Perl on Windows.
- CVE-2026-42250 (libbz2-1.0): the bug is in
bzip2recover, which is in thebzip2package, not in the image.
They are still counted as errors above. A reviewer should know that "other platform" answers are this tool's most frequent disagreement with a distro's VEX.
Cost and speed
- Test run: 912 findings in 1,372 s with 8 parallel calls on a busy shared gateway.
- Smoke (4 findings): 13,233 tokens, $0.0048 at the gateway list price. That is about 3,300 tokens and $0.0012 per model call, one call per finding that the rules do not settle.
- Hosted sample runs on the pre-release server: nginx, 12 findings, 14.3 s; python, 12 findings, 13.8 s.
Whole reports (no labels): how much noise the checks remove
In rules mode, with the same checks the model sees:
nginx:1.20.0: 168 of 550 Grype findings get a not_affected with strong evidence. 164 arevulnerable_code_not_in_execute_path, mostly libtiff, libwebp, libexpat, libxml2, libgd, libX11 and freetype, which only the image-filter and XSLT dynamic modules load, and the shipped nginx.conf loads neither. 4 arecomponent_not_present.python:3.10.0-slim-bullseye: 10 of 637.
No vendor labels exist for these images (Debian publishes no OpenVEX), so these counts are not accuracy.
Round trip into the scanner
The OpenVEX from the recorded nginx-1.20.0 run was fed back to Grype with
grype registry:nginx:1.20.0 --vex doc.openvex.json. Grype moved exactly the 4 not_affected findings (libwebp, libexpat
and two libtiff) to ignoredMatches and kept the other 546.
This needed one fix, found while testing. The image identity must be the repository digest, not Grype's manifest
digest, and the pkg:oci product must have no tag qualifier. Trivy's --vex was not tested.
Expected properties of the sample runs (for the rehearsal kit)
nginx-1.20.0, CVE-2023-4863 (libwebp6, CISA KEV):not_affected/vulnerable_code_not_in_execute_path. The loader chain is libwebp.so.6 <- libgd.so.3 <- ngx_http_image_filter_module.so, and no config file names that module.nginx-1.20.0-image-filter(the same image with aload_moduleline added): CVE-2023-4863 becomesaffected.- CVE-2022-0778 (libssl1.1 and openssl):
affectedin both nginx samples, because nginx links libssl. - CVE-2009-4487 (nginx): never
not_affected. NVD lists only 0.7.64, so the check is weak and the answer isunder_investigation. python-3.10.0-slim, CVE-2022-2097 (32-bit x86 only):not_affected/vulnerable_code_not_in_execute_path. CVE-2023-4911 (glibc):under_investigation, because the only evidence is weak.- Every model call has a signed receipt. The OpenVEX DSSE envelope verifies at
/vex/verifyand fails after any payload change. The evidence record verifies at/record/verify.
rehearsal/vex-triage/ checks 1, 3, 4 and 6 (10 checks, passed 10/10 on the pre-release server).
Honest limits
- Labels. They come from one vendor (Canonical) and two releases. Canonical's statements are package-level and include judgment calls, such as "I can't see how this could be exploited" and "disputed". This tool never makes those calls, which explains most of the low not_affected recall (32.5%). No Red Hat, Debian or Chainguard labels were used. Chainguard's advisories are CC BY-NC-ND, so they were left out of a commercial eval.
- Probes are constructed findings. A real CPE-matching scanner would flag some of them and not others, so the precision numbers depend on the mix. Read the per-class counts, not the ratio alone.
- The model adds little to the decisions. Rules only has the same precision and slightly higher agreement. The
model writes the impact statements and sends unclear cases to a person (
under_investigation), and the gates stop it from over-claiming: on the test set, 101 not_affected proposals without strong evidence became under_investigation. A buyer could runmode: "rules"(no GPU) and get about the same statuses. - Checks see the image as shipped. Configs mounted at run time and
dlopenby path are not seen. Library use is read from ELF links only. There is no call-graph reachability inside programs, and no language-level reachability (Python imports, Go symbols). - Calibration. It is fit on dev and separates little: the model says 95 on 308 of 320. Treat
probabilityas informational. - Distro packages only. Only distro packages are evaluated. Language packages (PyPI, npm, Go) use OSV ranges in the same code path but were not measured against labels.
- Who judged. The building agent chose the probe method and read the statements. There was no independent security engineer.