89 · Sales and marketing · Compliance and trust · live
Storefront accessibility pass
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted issues caught122 / 127test splitn = 12717 issue types on 30 pages of four test products; axe-core 52/52, keyboard and form passes 18/18, model 52/57. Misses: two wrong alts, three links named "Details".
- Clean items flagged0 / 207test splitn = 207Good alts, clear names and labelled fields on the same pages; advisories do not count.
- Clean pages with any finding0 / 7test splitn = 7
- Alt judgement: bad alts caught / good alts flagged111 / 113 and 0 / 38test splitn = 151Wrong product, colour swapped, file name, vague, keyword-stuffed or empty-but-needed; colour swaps 16/18.
- Blind comparison, alt cases right: Qwen3.8 / Opus 5.557 / 60 and 59 / 60test splitn = 60Claude Code Opus 5.5 as a blind sub-agent with the same instructions; Qwen's 3 misses are plain background photos it wants an alt for.
- Blind comparison, planted names caught: Qwen3.8 / Opus 5.510 / 12 and 11 / 12test splitn = 12Both left all 28 clean names alone.
- Dev: planted caught / clean items flagged73 / 74 and 0 / 119dev (tuned on)n = 193Prompts and empty-alt rules set here in three rounds.
Dataset
Harbor & Pine, a made-up shop: six invented products (2 dev, 4 test, split by product) plus home, cart and checkout pages; 50 pages with 201 planted issues of 17 types and 326 clean items; alt variants per image.
Caveats
- Same author built the shop, the plants, the checker and the labels; one synthetic template, no real themes.
- The demo scenarios use two test-split products and were run during development; the eval pages were not.
- Small n for some types: keyboard trap 2, unlinked error 2, label mismatch 1.
- The frontier comparison is one blind sample of 100 cases, scored by us.
- Latency was measured on a shared, loaded gateway.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 28 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 1.8 s
- Receipts
- 2
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- <$0.001
Self-host verification
Verified on 27 Sep 2026: Fresh clone of the branch into a clean directory, docker build of the api image with WITH_BROWSER=1, the api on host networking against the running local Qwen3.8-27B (vision, direct route); then torn down.
The rehearsal bundle passed 8 of 8 in 11 s; the three-page sample shop gave 13 findings (5 P1) and the clean shop none, receipts attested; a URL audit of a local http staging page (DECOSA_TESTRUNS_TARGETS + ALLOW_PRIVATE, 390 px) found its 6 issues with only GET requests reaching the server. The model server's own startup was not re-verified (no new GPU load).
Rehearsal bundle: storefront-accessibility-pass.zip (58 KB, 8 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (10.7, 10.9, 13.3 s) and 3 on a quiet gateway on 28 Sep (1.8, 1.8, 1.8 s). Cost is the median at list price, model calls included (range $0.0004 to $0.0004).
- Measured on one made-up shop template; real themes with lazy-loaded images, cookie banners and third-party widgets were not tested.
- The model asks for an alt on a product photo used as a background behind real text; the page audit reports that as an advisory, not a finding.
- Pages behind a login can only be audited self-hosted or pasted as HTML.
- The demo console audits the sample shop and pasted HTML; live URLs need an API key and a verified domain.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Looks at each image with its alt text (right, wrong, poor, needed but empty, decorative), reads offer text drawn into banners, judges link and button names in their context and the form's error messagesQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Opens each page in a private headless Chromium (one per audit), runs axe-core, the keyboard pass (Tab order, Enter on add-to-cart, visible focus, Escape from dialogs) and the empty-submit form pass; blocks every request that is not GET or HEADaxe-core 4.13.0 in headless Chromium (Playwright 1.58)MPL-2.0 (axe-core) + Apache-2.0 (Playwright) + BSD-3-Clause (Chromium)
- Merges the rule, keyboard, form and model results into findings ranked P1 (blocks a purchase) to P4, cites the WCAG 2.2 success criteria, writes a code fix per finding and seals the signed recorddecosa-api a11y module (decosa_api/verticals/a11y)AGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · rules, keyboard and forms, no model (CPU) (1)
- held-out test: planted issues found by the rule engine and the keyboard and form passes (no model): 70 of 70 of those types; the 57 model-layer issues are not checkeddecosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route
Standard · rules, keyboard, forms and the model (hosted demo) (2)
- held-out test (frozen): planted issues caught / clean items flagged / clean pages with any finding: 122 of 127 / 0 of 207 / 0 of 7decosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route
- blind comparison on 60 alt and 40 name cases: Qwen3.8-27B vs Claude Code Opus 5.5: alt 57/60 vs 59/60; planted names 10/12 vs 11/12decosa-api docs/evals/storefront-accessibility-pass.md, measured on our server 2026-09-27, gateway route