Eval: storefront accessibility pass (use case 89)
Run on our server on 27 Sep 2026, branch the pre-release branch, Qwen3.8-27B (NVFP4, vision) through the
model gateway (receipted), axe-core 4.13.0 in headless Chromium (Playwright 1.58). Runner: scripts/a11y_eval.py;
raw scores in docs/evals/storefront-accessibility-pass/.
Data
Harbor & Pine, a made-up shop (decosa_api/verticals/a11y/store.py). Six invented products: packshots drawn with
Pillow, lifestyle scenes drawn by Wan2.2-VACE-Fun-A14B (Apache-2.0) for use case 77, banners cut from those scenes with
offer text drawn in Lato (SIL OFL 1.1). No third-party store was crawled.
- Split by product: dev = Kettle & Crane tea and Brightloom candle, plus dev variants of the home, cart and checkout pages (20 pages); test = Fernleaf hand wash, Sol Fizz, Northwind D3 and Tidewell toothpaste, plus test variants of home, cart and checkout (30 pages). Link wordings, vague alt texts, file names and error wordings come from disjoint pools per split.
- Planted issues (17 types), each tagged on its element in a manifest the page never shows: img-no-alt,
alt-empty-informative, alt-wrong (another product's alt, or a colour swapped), alt-poor (file name, generic word,
keyword stuffing), image-text-missing (offer only in a banner), button-no-name, link-no-name, name-vague,
name-mismatch (visible "Add to cart", announced "Send"), input-no-label, low-contrast, html-no-lang,
kbd-not-focusable (add-to-cart or checkout as a
div), kbd-trap (dialog with no Escape), focus-invisible, error-unclear ("Validation failed (code 17)"), error-unlinked (error shown by a red border only). - Clean: every page also carries clean items (good alt, clear names, labelled fields), and each product and page type has a fully clean version.
- dev: 74 planted issues, 119 clean items, 5 clean pages. test: 127 planted, 207 clean items, 7 clean pages.
A planted issue is caught when a finding of a matching category lands on the same element (or field, or page for page-level issues). A clean item is flagged when any finding lands on it. Advisories (e.g. "weak name") do not count as findings either way.
Prompts and the empty-alt rules were set in three rounds on the dev split only (eval-round1, eval-round2, then the
final dev run). The test split was run once with everything frozen; nothing changed after it.
Results: the full audit on each page (rules + keyboard + forms + model)
| dev | test | |
|---|---|---|
| Planted issues caught | 73 / 74 | 122 / 127 |
| by the rule engine (axe-core) | 29 / 29 | 52 / 52 |
| by the keyboard and form passes | 10 / 10 | 18 / 18 |
| by the model | 34 / 35 | 52 / 57 |
| Clean items flagged | 0 / 119 | 0 / 207 |
| Clean pages with any finding | 0 / 5 | 0 / 7 |
| Model calls | 80 | 146 |
| Median seconds per page | 28.7 | 51.8 |
Test misses (5): two alt-wrong (a "pump bottle on a black marble counter" alt on a photo of the same bottle in another bath scene; a colour-swapped jar alt), three name-vague (links named "Details"; one came back as an advisory). Both "unattributed" findings per split are real: the close "×" on the planted keyboard-trap dialog is a span Tab never reaches.
Per test type: alt-empty-informative 11/11, alt-poor 18/18, alt-wrong 5/7, button-no-name 8/8, error-unclear 4/4, error-unlinked 2/2, focus-invisible 6/6, html-no-lang 8/8, image-text-missing 5/5, img-no-alt 12/12, input-no-label 9/9, kbd-not-focusable 8/8, kbd-trap 2/2, link-no-name 5/5, low-contrast 10/10, name-mismatch 1/1, name-vague 8/11.
What the model adds: with "judge": false (the Lite tier) the 57 model-layer test plants (wrong, poor and
empty-but-informative alt, vague and mismatched names, unclear errors, text only in images) are not checked at all;
axe-core passes every one of them because each has an alt, a name or a message.
Results: the alt-text judgement alone
Every image x 8 alt variants (good, short, another product's alt, colour swapped, file name, vague, keyword-stuffed, empty), plus each banner with and without drawn text.
| dev | test | |
|---|---|---|
| Bad alts caught (wrong, poor or needed-but-empty) | 44 / 48 | 111 / 113 |
| Good alts flagged | 0 / 16 | 0 / 38 |
| Colour-swapped alt caught | 4 / 8 | 16 / 18 |
| Banner with drawn offer: flagged, offer read back | 2 / 2 | 4 / 4 |
| Plain banner photo with alt="" judged decorative | 0 / 2 | 0 / 4 |
The plain-banner row is a disagreement over labels as much as a miss: the photo behind the banner text is a product scene, and Qwen asks for an alt. In the page audit this case is an advisory ("background image"), not a finding, which is why the clean pages stay clean.
Blind frontier comparison
Judge: Claude Code Opus 5.5 as a blind sub-agent, on a random sample from the test split (60 alt cases, 40 link and button names, 12 of them planted). It saw only the images, the alt and context, the controls, and the same instructions and output schema as Qwen; no labels or planted notes. Its alt output went through the same parser as Qwen's (the facts about missing or empty alt are set in code). Qwen was re-run fresh on the same cases.
| Qwen3.8-27B | Opus 5.5 (blind) | |
|---|---|---|
| Alt cases right | 57 / 60 | 59 / 60 |
| of which plain banners (alt="" is right) | 0 / 3 | 3 / 3 |
| of which good alts left alone | 7 / 7 | 6 / 7 |
| Planted vague or mismatched names caught | 10 / 12 | 11 / 12 |
| Clean names left alone | 28 / 28 | 28 / 28 |
The gap is the decorative-background judgement (3 cases) and one vague name. On wrong, poor, stuffed, file-name and empty alts both are 100% on this sample.
Cost and speed
Test split: 146 model calls, 119,442 prompt and 17,765 generated tokens for 30 pages, i.e. about 4,000 prompt and 590 generated tokens per page: about $0.0021 per page at the gateway list price ($0.30 / $1.50 per million). Latency was measured with the alt eval running in parallel on a shared gateway: median 51.8 s per page on test, 28.7 s on dev. The hosted demo's three-page shop took 89 to 149 s under that load; the same audit on the self-hosted check (direct route, lighter load) took 10.7 to 12.6 s. Smoke run (one tiny page, 2 model calls): about $0.0004, 33.6 to 36.3 s under load.
Checkable properties of the demo (rehearsal bundle rehearsal/storefront-accessibility-pass/)
- Add to cart as a
divis found by the keyboard pass and ranked P1 (WCAG 2.1.1). - The unlabelled quantity field is found by axe-core (WCAG 1.3.1 / 4.1.2 / 3.3.2).
- The packshot alt that calls the orange can purple is flagged
alt-wrongby the model. - The offer that exists only inside the banner is flagged
image-text(WCAG 1.4.5 / 1.1.1). - Every model call has a receipt, and the signed record verifies at
/record/verify; a changed priority fails. - The clean sample shop, checked without the model, has no findings.
Caveats
- Same author built the shop, the plants, the checker and the labels; synthetic pages, one shop template. Real themes (Shopify, WooCommerce) have more markup, lazy-loaded images, cookie banners and third-party widgets; not measured.
- The demo scenarios use test-split products (Sol Fizz, Fernleaf) and were run during development, though the eval pages and their variants were not.
- Small n for some types (kbd-trap 2, error-unlinked 2, name-mismatch 1).
- The frontier judge is a single blind sample (n = 100) scored by us against our labels.
- What it does not check at all: screen-reader output, reading order, zoom and reflow, captions, cognitive load, every widget, pages behind a login on the hosted service.
- axe-core's own detection rate is not measured here beyond the planted rule-type issues.