Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: storefront accessibility pass (use case 89)

Run on our server on 27 Sep 2026, branch the pre-release branch, Qwen3.8-27B (NVFP4, vision) through the model gateway (receipted), axe-core 4.13.0 in headless Chromium (Playwright 1.58). Runner: scripts/a11y_eval.py; raw scores in docs/evals/storefront-accessibility-pass/.

Data

Harbor & Pine, a made-up shop (decosa_api/verticals/a11y/store.py). Six invented products: packshots drawn with Pillow, lifestyle scenes drawn by Wan2.2-VACE-Fun-A14B (Apache-2.0) for use case 77, banners cut from those scenes with offer text drawn in Lato (SIL OFL 1.1). No third-party store was crawled.

  • Split by product: dev = Kettle & Crane tea and Brightloom candle, plus dev variants of the home, cart and checkout pages (20 pages); test = Fernleaf hand wash, Sol Fizz, Northwind D3 and Tidewell toothpaste, plus test variants of home, cart and checkout (30 pages). Link wordings, vague alt texts, file names and error wordings come from disjoint pools per split.
  • Planted issues (17 types), each tagged on its element in a manifest the page never shows: img-no-alt, alt-empty-informative, alt-wrong (another product's alt, or a colour swapped), alt-poor (file name, generic word, keyword stuffing), image-text-missing (offer only in a banner), button-no-name, link-no-name, name-vague, name-mismatch (visible "Add to cart", announced "Send"), input-no-label, low-contrast, html-no-lang, kbd-not-focusable (add-to-cart or checkout as a div), kbd-trap (dialog with no Escape), focus-invisible, error-unclear ("Validation failed (code 17)"), error-unlinked (error shown by a red border only).
  • Clean: every page also carries clean items (good alt, clear names, labelled fields), and each product and page type has a fully clean version.
  • dev: 74 planted issues, 119 clean items, 5 clean pages. test: 127 planted, 207 clean items, 7 clean pages.

A planted issue is caught when a finding of a matching category lands on the same element (or field, or page for page-level issues). A clean item is flagged when any finding lands on it. Advisories (e.g. "weak name") do not count as findings either way.

Prompts and the empty-alt rules were set in three rounds on the dev split only (eval-round1, eval-round2, then the final dev run). The test split was run once with everything frozen; nothing changed after it.

Results: the full audit on each page (rules + keyboard + forms + model)

dev test
Planted issues caught 73 / 74 122 / 127
by the rule engine (axe-core) 29 / 29 52 / 52
by the keyboard and form passes 10 / 10 18 / 18
by the model 34 / 35 52 / 57
Clean items flagged 0 / 119 0 / 207
Clean pages with any finding 0 / 5 0 / 7
Model calls 80 146
Median seconds per page 28.7 51.8

Test misses (5): two alt-wrong (a "pump bottle on a black marble counter" alt on a photo of the same bottle in another bath scene; a colour-swapped jar alt), three name-vague (links named "Details"; one came back as an advisory). Both "unattributed" findings per split are real: the close "×" on the planted keyboard-trap dialog is a span Tab never reaches.

Per test type: alt-empty-informative 11/11, alt-poor 18/18, alt-wrong 5/7, button-no-name 8/8, error-unclear 4/4, error-unlinked 2/2, focus-invisible 6/6, html-no-lang 8/8, image-text-missing 5/5, img-no-alt 12/12, input-no-label 9/9, kbd-not-focusable 8/8, kbd-trap 2/2, link-no-name 5/5, low-contrast 10/10, name-mismatch 1/1, name-vague 8/11.

What the model adds: with "judge": false (the Lite tier) the 57 model-layer test plants (wrong, poor and empty-but-informative alt, vague and mismatched names, unclear errors, text only in images) are not checked at all; axe-core passes every one of them because each has an alt, a name or a message.

Results: the alt-text judgement alone

Every image x 8 alt variants (good, short, another product's alt, colour swapped, file name, vague, keyword-stuffed, empty), plus each banner with and without drawn text.

dev test
Bad alts caught (wrong, poor or needed-but-empty) 44 / 48 111 / 113
Good alts flagged 0 / 16 0 / 38
Colour-swapped alt caught 4 / 8 16 / 18
Banner with drawn offer: flagged, offer read back 2 / 2 4 / 4
Plain banner photo with alt="" judged decorative 0 / 2 0 / 4

The plain-banner row is a disagreement over labels as much as a miss: the photo behind the banner text is a product scene, and Qwen asks for an alt. In the page audit this case is an advisory ("background image"), not a finding, which is why the clean pages stay clean.

Blind frontier comparison

Judge: Claude Code Opus 5.5 as a blind sub-agent, on a random sample from the test split (60 alt cases, 40 link and button names, 12 of them planted). It saw only the images, the alt and context, the controls, and the same instructions and output schema as Qwen; no labels or planted notes. Its alt output went through the same parser as Qwen's (the facts about missing or empty alt are set in code). Qwen was re-run fresh on the same cases.

Qwen3.8-27B Opus 5.5 (blind)
Alt cases right 57 / 60 59 / 60
of which plain banners (alt="" is right) 0 / 3 3 / 3
of which good alts left alone 7 / 7 6 / 7
Planted vague or mismatched names caught 10 / 12 11 / 12
Clean names left alone 28 / 28 28 / 28

The gap is the decorative-background judgement (3 cases) and one vague name. On wrong, poor, stuffed, file-name and empty alts both are 100% on this sample.

Cost and speed

Test split: 146 model calls, 119,442 prompt and 17,765 generated tokens for 30 pages, i.e. about 4,000 prompt and 590 generated tokens per page: about $0.0021 per page at the gateway list price ($0.30 / $1.50 per million). Latency was measured with the alt eval running in parallel on a shared gateway: median 51.8 s per page on test, 28.7 s on dev. The hosted demo's three-page shop took 89 to 149 s under that load; the same audit on the self-hosted check (direct route, lighter load) took 10.7 to 12.6 s. Smoke run (one tiny page, 2 model calls): about $0.0004, 33.6 to 36.3 s under load.

Checkable properties of the demo (rehearsal bundle rehearsal/storefront-accessibility-pass/)

  1. Add to cart as a div is found by the keyboard pass and ranked P1 (WCAG 2.1.1).
  2. The unlabelled quantity field is found by axe-core (WCAG 1.3.1 / 4.1.2 / 3.3.2).
  3. The packshot alt that calls the orange can purple is flagged alt-wrong by the model.
  4. The offer that exists only inside the banner is flagged image-text (WCAG 1.4.5 / 1.1.1).
  5. Every model call has a receipt, and the signed record verifies at /record/verify; a changed priority fails.
  6. The clean sample shop, checked without the model, has no findings.

Caveats

  • Same author built the shop, the plants, the checker and the labels; synthetic pages, one shop template. Real themes (Shopify, WooCommerce) have more markup, lazy-loaded images, cookie banners and third-party widgets; not measured.
  • The demo scenarios use test-split products (Sol Fizz, Fernleaf) and were run during development, though the eval pages and their variants were not.
  • Small n for some types (kbd-trap 2, error-unlinked 2, name-mismatch 1).
  • The frontier judge is a single blind sample (n = 100) scored by us against our labels.
  • What it does not check at all: screen-reader output, reading order, zoom and reflow, captions, cognitive load, every widget, pages behind a login on the hosted service.
  • axe-core's own detection rate is not measured here beyond the planted rule-type issues.