Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: the site kit's JSON-LD validator (use case 181 and the Site checks block)

29 Sep 2026, port-elmo-opus. Script: scripts/site_kit_eval.py. Cases and results: docs/evals/site-kit/.

What was measured

  1. Planted errors. Every JSON-LD example in schemaorg/schemaorg (Apache-2.0, commit 4acb22f) whose top-level type is in a family the validator has rules for: 55 unique objects (the families: Article 2, BlogPosting 1, BreadcrumbList 1, LocalBusiness 24, NewsArticle 1, Organization 15, Product 10, WebSite 1; there is no FAQPage example). Code planted one error per variant, two variants per example, each with the path and severity it must produce: a required property removed or emptied, a placeholder value, a wrong @context, no @type, a nested error (an offer without a price, a breadcrumb item without a position). A planted error counts as caught only when the expected issue appears at its path and was not already there on the clean example.
  2. Parity. The TypeScript kit and the Python mirror on the same 165 cases, and on the homepages of 29 public sites (the HTML was fetched once; only derived results are kept, in real.json).

Results

Result
Planted errors caught 109 of 110 (bad-context 18/18, empty-required 22/22, nested 2/2, no-type 25/25, placeholder 21/22, remove-required 21/21)
TypeScript and Python give the same result 165 of 165 cases; 29 of 29 real homepages
Clean schema.org examples with at least one error 11 of 55
Clean examples with at least one warning 47 of 55 (mostly example.com placeholders, which the examples use on purpose)
  • The one missed plant is a scoring artifact: the example's url was already example.com, so planting a placeholder added no new warning.
  • The errors on clean examples are mostly real by the validator's rules, which follow what Google needs for a rich result: schema.org's examples are partial illustrations (an Article with only a name, a Product without offers). One is a false error: a BreadcrumbList that puts each item's name inside item. It was left as found; the rules were not changed after the cases were built.
  • Real homepages: 10 of 29 had JSON-LD; none had an unreadable block; nine answered 403 to a non-browser client.

Site check timing

POST /geo/site-check on 24 public sites from the pre-release server: see content/verticals/geo-audit/stack.json verification (site_check) on the site for the measured p50 and p95. No model call, so the cost is CPU time only.

Also changed by this work

The geo crawl now reads robots.txt by RFC 9309 instead of Python's urllib.robotparser, which takes the first matching rule in file order, ignores * and $, and matches crawler names by substring ("User-agent: Claude" blocked Claude-SearchBot). tests/test_reviewreply.py::test_site_check_route keeps that from coming back.

Limits

  • Rules for seven schema.org families only; other types get the @context and placeholder checks.
  • It reads the HTML a crawler without JavaScript sees.
  • The planted errors were written by the same author as the rules.