Skip to content
decosa

Developers · Building blocks

Site checks

Deterministic checks for whether a website is ready for AI answers, and a guard for public review replies. robots.txt read the way the AI crawlers document it, llms.txt, every JSON-LD object checked property by property, and a reply checker that never lets a health or care practice confirm a patient. No model in the checks themselves, open source, and the same code in TypeScript and Python.

Measured 2026-09-29. Site check eval · Review reply eval.

Watch a real run

Loading the tool…

How it works

  1. robots.txt is parsed by RFC 9309: every group naming a crawler is combined, the longest matching rule wins and allow wins a tie, with * and $ supported. Python's standard robots parser takes the first match in file order and matches crawler names by substring, so a group for "Claude" would wrongly block Claude-SearchBot.
  2. JSON-LD is read from the raw HTML before any script is stripped, including arrays, @graph and commented or entity-encoded blocks. Each object is checked against its family's rules (LocalBusiness and its subtypes, Organization, Article, Product, FAQPage, BreadcrumbList, WebSite) and for placeholder values that must never ship.
  3. llms.txt is checked for a title line, a summary and links, and flagged when a site serves an HTML page in its place.
  4. The review guard checks a reply for offers, contact details or numbers nobody gave, placeholders and arguing; for health and care businesses, for anything that confirms a patient, a visit, a treatment, a wait, a bill, a family member or a name.
  5. Shared fixtures run in both languages' test suites (vitest and pytest), so the TypeScript kit and the Python mirror return byte-identical results.

Speed and cost

Site check, one site (24 public sites)p50 1.7 s, p95 4.3 s; no model call
Review reply, one reply (100 replies, shared gateway)p50 1.8 s, p95 4.2 s; about $0.000441 a reply at list price
Checking a reply you wrote (POST /reviews/check)code only, no model call

Results

Site check: Planted JSON-LD errors caught109 / 110 (schema.org's own examples with one error planted by code; by kind: bad-context 18/18, empty-required 22/22, nested 2/2, no-type 25/25, placeholder 21/22, remove-required 21/21)
Site check: Same result in TypeScript and Python165 / 165 (and 29 of 29 real homepages)
Site check: schema.org examples flagged with an error11 / 55 (mostly partial examples missing what Google needs for a rich result; one is a false error (a breadcrumb name inside item))
Review reply: Healthcare replies that confirm a patient (blind judge)0 / 40 (held-out set generated after the guard was frozen; target 0)
Review reply: Replies a business could post as written (blind judge)93 / 100 (healthcare 38/40, other 55/60)
Review reply: Healthcare replies that fell back to the fixed safe reply24 / 40 (the price of blocking broadly)
Review reply: Replies with an offer or a fact nobody gave (blind judge)4 / 100

API and library

POST /geo/site-check{domain, name?} -> robots access per AI crawler, llms.txt, every JSON-LD object's errors and warnings, page text, fixes with snippets and a signed record. No model call. Token or dk_ key for geo-audit.
POST /reviews/check{review, rating, business_name, business_type?, reply} -> ok, problems, the blocked phrases. No model call. Token or dk_ key for review-reply.
POST /reviews/replyThe drafting job on top of the guard: a draft, a redraft naming the blocked words, else a fixed safe reply; receipted model calls.
@decosa/site-kit (TypeScript)checkPageJsonLd, validateSchemaLocally, robotsAccess, llmsTxtCheck, checkLiveMeta, hasMeaningfulAltText, checkReply, isHealthcareBusiness, generateSchemaMarkup. In clients/site-kit of the decosa-api source; not on npm yet.
decosa-api (Python)decosa_api.verticals.geo.schema_check, geo.site_rules and reviewreply.guard, Apache-2.0 exceptions to the AGPL.

Where it is used

  • Check your site is ready for AI answersbuilt

    The site check, with the fix to paste for each finding.

  • Reply to a review without breaking patient privacybuilt

    The review guard with a drafting model and a privacy check.

  • ElmoSEOswitching

    ElmoSEO's validator, schema generator, live-page check, alt-text rule and review guard came from here and move to importing this kit (on a branch, before review).

  • Verify a fix is liveunblocked

    checkLiveMeta and checkLiveImageAlt confirm a CMS change really reached the public page, through CDN caches.

Licences

@decosa/site-kit and its Python mirrorsApache-2.0, first written for ElmoSEO and relicensed by its author
schema.org examples used in the evalApache-2.0 (schemaorg/schemaorg)
The review-reply job and the geo crawlAGPL-3.0-or-later, like the rest of decosa-api

What it does not do

  • It doesn't render JavaScript: it reports what a crawler without JavaScript sees, which is what most AI crawlers see.
  • It has rules for seven schema.org families; other types get only the context and placeholder checks.
  • It doesn't measure whether answer engines mention you: that is the answer panel's job.
  • The review guard is English only, and a new phrasing can still slip through: read every reply before posting.
All building blocks