Eval: the site kit's JSON-LD validator (use case 181 and the Site checks block)
29 Sep 2026, port-elmo-opus. Script: scripts/site_kit_eval.py. Cases and results: docs/evals/site-kit/.
What was measured
- Planted errors. Every JSON-LD example in schemaorg/schemaorg (Apache-2.0, commit
4acb22f) whose top-level type is in a family the validator has rules for: 55 unique objects (the families: Article 2, BlogPosting 1, BreadcrumbList 1, LocalBusiness 24, NewsArticle 1, Organization 15, Product 10, WebSite 1; there is no FAQPage example). Code planted one error per variant, two variants per example, each with the path and severity it must produce: a required property removed or emptied, a placeholder value, a wrong@context, no@type, a nested error (an offer without a price, a breadcrumb item without a position). A planted error counts as caught only when the expected issue appears at its path and was not already there on the clean example. - Parity. The TypeScript kit and the Python mirror on the same 165 cases, and on the homepages of
29 public sites (the HTML was fetched once; only derived results are kept, in
real.json).
Results
| Result | |
|---|---|
| Planted errors caught | 109 of 110 (bad-context 18/18, empty-required 22/22, nested 2/2, no-type 25/25, placeholder 21/22, remove-required 21/21) |
| TypeScript and Python give the same result | 165 of 165 cases; 29 of 29 real homepages |
| Clean schema.org examples with at least one error | 11 of 55 |
| Clean examples with at least one warning | 47 of 55 (mostly example.com placeholders, which the examples use on purpose) |
- The one missed plant is a scoring artifact: the example's
urlwas alreadyexample.com, so planting a placeholder added no new warning. - The errors on clean examples are mostly real by the validator's rules, which follow what Google needs for a rich result:
schema.org's examples are partial illustrations (an Article with only a name, a Product without offers). One is a false
error: a BreadcrumbList that puts each item's name inside
item. It was left as found; the rules were not changed after the cases were built. - Real homepages: 10 of 29 had JSON-LD; none had an unreadable block; nine answered 403 to a non-browser client.
Site check timing
POST /geo/site-check on 24 public sites from the pre-release server: see content/verticals/geo-audit/stack.json
verification (site_check) on the site for the measured p50 and p95. No model call, so the cost is CPU time only.
Also changed by this work
The geo crawl now reads robots.txt by RFC 9309 instead of Python's urllib.robotparser, which takes the first matching
rule in file order, ignores * and $, and matches crawler names by substring ("User-agent: Claude" blocked
Claude-SearchBot). tests/test_reviewreply.py::test_site_check_route keeps that from coming back.
Limits
- Rules for seven schema.org families only; other types get the
@contextand placeholder checks. - It reads the HTML a crawler without JavaScript sees.
- The planted errors were written by the same author as the rules.