Site checks
Deterministic checks for whether a website is ready for AI answers, and a guard for public review replies. robots.txt read the way the AI crawlers document it, llms.txt, every JSON-LD object checked property by property, and a reply checker that never lets a health or care practice confirm a patient. No model in the checks themselves, open source, and the same code in TypeScript and Python.
Measured 2026-09-29. Site check eval · Review reply eval.
Watch a real run
Loading the tool…
How it works
- robots.txt is parsed by RFC 9309: every group naming a crawler is combined, the longest matching rule wins and allow wins a tie, with * and $ supported. Python's standard robots parser takes the first match in file order and matches crawler names by substring, so a group for "Claude" would wrongly block Claude-SearchBot.
- JSON-LD is read from the raw HTML before any script is stripped, including arrays, @graph and commented or entity-encoded blocks. Each object is checked against its family's rules (LocalBusiness and its subtypes, Organization, Article, Product, FAQPage, BreadcrumbList, WebSite) and for placeholder values that must never ship.
- llms.txt is checked for a title line, a summary and links, and flagged when a site serves an HTML page in its place.
- The review guard checks a reply for offers, contact details or numbers nobody gave, placeholders and arguing; for health and care businesses, for anything that confirms a patient, a visit, a treatment, a wait, a bill, a family member or a name.
- Shared fixtures run in both languages' test suites (vitest and pytest), so the TypeScript kit and the Python mirror return byte-identical results.
Speed and cost
| Site check, one site (24 public sites) | p50 1.7 s, p95 4.3 s; no model call |
| Review reply, one reply (100 replies, shared gateway) | p50 1.8 s, p95 4.2 s; about $0.000441 a reply at list price |
| Checking a reply you wrote (POST /reviews/check) | code only, no model call |
Results
| Site check: Planted JSON-LD errors caught | 109 / 110 (schema.org's own examples with one error planted by code; by kind: bad-context 18/18, empty-required 22/22, nested 2/2, no-type 25/25, placeholder 21/22, remove-required 21/21) |
| Site check: Same result in TypeScript and Python | 165 / 165 (and 29 of 29 real homepages) |
| Site check: schema.org examples flagged with an error | 11 / 55 (mostly partial examples missing what Google needs for a rich result; one is a false error (a breadcrumb name inside item)) |
| Review reply: Healthcare replies that confirm a patient (blind judge) | 0 / 40 (held-out set generated after the guard was frozen; target 0) |
| Review reply: Replies a business could post as written (blind judge) | 93 / 100 (healthcare 38/40, other 55/60) |
| Review reply: Healthcare replies that fell back to the fixed safe reply | 24 / 40 (the price of blocking broadly) |
| Review reply: Replies with an offer or a fact nobody gave (blind judge) | 4 / 100 |
API and library
| POST /geo/site-check | {domain, name?} -> robots access per AI crawler, llms.txt, every JSON-LD object's errors and warnings, page text, fixes with snippets and a signed record. No model call. Token or dk_ key for geo-audit. |
| POST /reviews/check | {review, rating, business_name, business_type?, reply} -> ok, problems, the blocked phrases. No model call. Token or dk_ key for review-reply. |
| POST /reviews/reply | The drafting job on top of the guard: a draft, a redraft naming the blocked words, else a fixed safe reply; receipted model calls. |
| @decosa/site-kit (TypeScript) | checkPageJsonLd, validateSchemaLocally, robotsAccess, llmsTxtCheck, checkLiveMeta, hasMeaningfulAltText, checkReply, isHealthcareBusiness, generateSchemaMarkup. In clients/site-kit of the decosa-api source; not on npm yet. |
| decosa-api (Python) | decosa_api.verticals.geo.schema_check, geo.site_rules and reviewreply.guard, Apache-2.0 exceptions to the AGPL. |
Where it is used
Check your site is ready for AI answersbuilt
The site check, with the fix to paste for each finding.
Reply to a review without breaking patient privacybuilt
The review guard with a drafting model and a privacy check.
ElmoSEOswitching
ElmoSEO's validator, schema generator, live-page check, alt-text rule and review guard came from here and move to importing this kit (on a branch, before review).
Verify a fix is liveunblocked
checkLiveMeta and checkLiveImageAlt confirm a CMS change really reached the public page, through CDN caches.
Licences
| @decosa/site-kit and its Python mirrors | Apache-2.0, first written for ElmoSEO and relicensed by its author |
| schema.org examples used in the eval | Apache-2.0 (schemaorg/schemaorg) |
| The review-reply job and the geo crawl | AGPL-3.0-or-later, like the rest of decosa-api |
What it does not do
- It doesn't render JavaScript: it reports what a crawler without JavaScript sees, which is what most AI crawlers see.
- It has rules for seven schema.org families; other types get only the context and placeholder checks.
- It doesn't measure whether answer engines mention you: that is the answer panel's job.
- The review guard is English only, and a new phrasing can still slip through: read every reply before posting.