Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: patent claim-support checker (36)

Run on our server, 25 Sep 2026, against the pre-release server (127.0.0.1:8436) on the hosted route: Qwen3.8-27B NVFP4 through the model gateway, temperature 0, one receipted call per claim element, the whole specification in the prompt (all eight eval specifications are under 48,000 characters). Script: scripts/patent_eval.py; data and raw results: docs/evals/patent-claim-support/. Metrics: metrics.json. Nothing was spent beyond the gateway's own metering (about $0.32 of list-price tokens per full support run).

Data

  • Eight granted US patents (public domain; text from the USPTO via the Google Patents public page, paragraph numbers as filed): 2 dev (US11043724B2 filter, US12078227B2 stroke device), 6 test (US10000000B2 LADAR, US11258908B2 headphone audio, US11456358B2 graphene, US11680592B2 structural panels, US12043097B2 roof drive, US12174590B2 watch). 212 claim elements in all (46 dev, 166 test once list headings are left out).
  • Labels: for every element, supported / partial / unsupported, every paragraph that describes it and the 1-3 a practitioner would cite, and whether the call is clear or a judgment call. Written by Claude (two separate agents, each reading the full specification), not by a registered practitioner. We did not have examiner reasons for allowance to check against; granted claims were examined, so most elements should be supported.
  • Antecedent set: 213 more granted patents drawn at random (US 11,000,000-12,400,000, B2), claims only, in four rounds (43, 55, 60, 55 patents; 3,478 claims).

1. Support map (element → paragraphs)

Dev (2 patents) Test (6 patents)
Elements scored 46 166
Supported-or-not call agrees with the labels 97.8% 93.4%
... on the elements labelled as clear calls 100% (41) 100% (146)
Supported elements: a cited paragraph is one the labels list 100% 100% (150)
Share of cited paragraphs the labels list 96.9% 93.8%
Elements labelled partly supported that the model flags (recall) 3 of 3 5 of 13
Model flags that match a label (precision) 3 of 4 5 of 8
  • What it is good at: finding where an element is described. On every supported test element at least one cited paragraph is one the labeller listed.
  • What it misses: partial support. All 8 misses and all 3 false flags on the test set are elements the labeller marked as judgment calls: a limitation described only in a different embodiment (the watch's battery size, claims 3-4), an arrangement the specification only implies (LADAR claims 10 and 20), an order of steps stated explicitly for one embodiment only (panels claim 1). Treat "supported" as "here is where to look", not as "no 112(a) problem".
  • Repeatability: three full runs (the second and third after a parser change that stopped judging list headings such as "the laser source further comprises:", which removed three false flags) gave test agreement of 93.6%, 91.0% and 93.4%. Elements near the line flip between weak and none or weak and strong between runs; the prompt was not changed after the first test run.
  • Cost and time: 166 calls, 996k prompt tokens (about 6,000 per element, mostly the specification, served from the model server's prefix cache) and 11k generated: $0.053 per application at the gateway's list price ($0.30 / $1.50 per million). 15-88 s per whole patent on the shared gateway (17-44 elements, 6 calls in flight). The 13-element demo sample takes 5.9-9.2 s hosted when the gateway is quiet (36 s once under other load) and 3.4 s self-hosted on the direct route.

2. Planted defects (6 test patents, final run)

plant() makes five changes per patent and the check runs on the result:

Defect Found
Every paragraph that describes one element deleted (the labelled ones plus every paragraph naming its main thing) 5 of 6
An introduction "a X" in claim 1 turned into "the X" 6 of 6
A reference in a dependent claim given a modifier with no basis ("the X" → "the second X") 5 of 5
A dependent claim pointed at a claim that does not exist 6 of 6
A dependent claim pointed at a later claim 6 of 6
A claim noun renamed everywhere to a word the specification never uses 6 of 6

The support-removal miss (panels, element 1.2) is a planting gap: paragraph [0035], left in, still describes forming a recess behind the mount, and the model cited it. An earlier run missed a different patent the same way (headphone, where [0005] and [0030] still describe the blending) and caught the panels one; over two runs 10 of 12 removals were flagged, and both misses cite a paragraph that still describes the element.

3. False alarms on clean granted claims

  • Code checks on the 6 test patents as granted: 15 issues. By our reading 12 are genuine: "the plurality of sample components" after "a plurality of components" (LADAR claims 1, 9, 11), "the atmosphere pressure chemical vapor deposition furnace" never introduced (graphene claims 1, 11), "the graphene film" after "a graphene layer" (claims 8, 18), "the external rotor wheel" never introduced (roof drive claims 5-7), and the claim words "discrete" and "outer" that the specifications never use. 3 are false: "the structure" of the metal (inherent, twice) and "the heating step" after "selectively heating" (implicit).
  • Support map on the same patents: 3 elements flagged that the labels call supported (all judgment calls, see 1).

4. Antecedent basis on random granted claims (code only)

Each round's flags were judged real or false (real = no explicit antecedent that a careful examiner could object to under MPEP 2173.05(e), arguable cases included; false = implicit or inherent antecedent, idiom, acronym, formula variable, superlative, spelling slip, or a phrase the extractor ran into a verb). The rules were then improved on that round, and the next round was drawn fresh; round 4 was drawn after the last change and is the held-out number. Judges: Claude for rounds 1-2, separate Claude agents given a written rubric for rounds 3-4.

Round Patents Claims Flags when judged Genuine Precision when judged Precision with the final rules
1 (dev) 43 730 82 31 38% 79% (30 of 31 genuine kept)
2 (dev) 55 920 113 52 46% 81% (50 of 52)
3 (dev) 60 937 203 58 29% 66% (56 of 58)
4 (held out) 55 891 66 36 55% 55%
  • On unseen granted claims, about one flag in two is a genuine slip, at about 1.2 flags per patent. The rest are mostly references a practitioner would accept (a verb form, an inherent property, an idiom) and phrase-extraction errors. The check is lenient by design; it is a proofreading aid, and each flag shows the closest earlier term.
  • Recall is measured only on planted errors (11 of 11 above); we did not label every reference in the random claims.
  • Each round surfaced new false-alarm classes, so expect precision on other art units to differ.

Not measured

  • Labels by a registered practitioner, or examiner reasons for allowance.
  • Long specifications (over 48,000 characters) in retrieved mode; the eight eval patents all went in whole.
  • The 32 GB card and the wanted GLM-5.3-Flash tier.
  • Drafting help, reference numerals and figures (not built).