Eval: patent claim-support checker (36)
Run on our server, 25 Sep 2026, against the pre-release server (127.0.0.1:8436) on the hosted route: Qwen3.8-27B NVFP4 through
the model gateway, temperature 0, one receipted call per claim element, the whole specification in the prompt (all
eight eval specifications are under 48,000 characters). Script: scripts/patent_eval.py; data and raw results:
docs/evals/patent-claim-support/. Metrics: metrics.json. Nothing was spent beyond the gateway's own metering
(about $0.32 of list-price tokens per full support run).
Data
- Eight granted US patents (public domain; text from the USPTO via the Google Patents public page, paragraph
numbers as filed): 2 dev (
US11043724B2filter,US12078227B2stroke device), 6 test (US10000000B2LADAR,US11258908B2headphone audio,US11456358B2graphene,US11680592B2structural panels,US12043097B2roof drive,US12174590B2watch). 212 claim elements in all (46 dev, 166 test once list headings are left out). - Labels: for every element, supported / partial / unsupported, every paragraph that describes it and the 1-3 a practitioner would cite, and whether the call is clear or a judgment call. Written by Claude (two separate agents, each reading the full specification), not by a registered practitioner. We did not have examiner reasons for allowance to check against; granted claims were examined, so most elements should be supported.
- Antecedent set: 213 more granted patents drawn at random (US 11,000,000-12,400,000, B2), claims only, in four rounds (43, 55, 60, 55 patents; 3,478 claims).
1. Support map (element → paragraphs)
| Dev (2 patents) | Test (6 patents) | |
|---|---|---|
| Elements scored | 46 | 166 |
| Supported-or-not call agrees with the labels | 97.8% | 93.4% |
| ... on the elements labelled as clear calls | 100% (41) | 100% (146) |
| Supported elements: a cited paragraph is one the labels list | 100% | 100% (150) |
| Share of cited paragraphs the labels list | 96.9% | 93.8% |
| Elements labelled partly supported that the model flags (recall) | 3 of 3 | 5 of 13 |
| Model flags that match a label (precision) | 3 of 4 | 5 of 8 |
- What it is good at: finding where an element is described. On every supported test element at least one cited paragraph is one the labeller listed.
- What it misses: partial support. All 8 misses and all 3 false flags on the test set are elements the labeller marked as judgment calls: a limitation described only in a different embodiment (the watch's battery size, claims 3-4), an arrangement the specification only implies (LADAR claims 10 and 20), an order of steps stated explicitly for one embodiment only (panels claim 1). Treat "supported" as "here is where to look", not as "no 112(a) problem".
- Repeatability: three full runs (the second and third after a parser change that stopped judging list headings
such as "the laser source further comprises:", which removed three false flags) gave test agreement of 93.6%, 91.0%
and 93.4%. Elements near the line flip between
weakandnoneorweakandstrongbetween runs; the prompt was not changed after the first test run. - Cost and time: 166 calls, 996k prompt tokens (about 6,000 per element, mostly the specification, served from the model server's prefix cache) and 11k generated: $0.053 per application at the gateway's list price ($0.30 / $1.50 per million). 15-88 s per whole patent on the shared gateway (17-44 elements, 6 calls in flight). The 13-element demo sample takes 5.9-9.2 s hosted when the gateway is quiet (36 s once under other load) and 3.4 s self-hosted on the direct route.
2. Planted defects (6 test patents, final run)
plant() makes five changes per patent and the check runs on the result:
| Defect | Found |
|---|---|
| Every paragraph that describes one element deleted (the labelled ones plus every paragraph naming its main thing) | 5 of 6 |
| An introduction "a X" in claim 1 turned into "the X" | 6 of 6 |
| A reference in a dependent claim given a modifier with no basis ("the X" → "the second X") | 5 of 5 |
| A dependent claim pointed at a claim that does not exist | 6 of 6 |
| A dependent claim pointed at a later claim | 6 of 6 |
| A claim noun renamed everywhere to a word the specification never uses | 6 of 6 |
The support-removal miss (panels, element 1.2) is a planting gap: paragraph [0035], left in, still describes forming a recess behind the mount, and the model cited it. An earlier run missed a different patent the same way (headphone, where [0005] and [0030] still describe the blending) and caught the panels one; over two runs 10 of 12 removals were flagged, and both misses cite a paragraph that still describes the element.
3. False alarms on clean granted claims
- Code checks on the 6 test patents as granted: 15 issues. By our reading 12 are genuine: "the plurality of sample components" after "a plurality of components" (LADAR claims 1, 9, 11), "the atmosphere pressure chemical vapor deposition furnace" never introduced (graphene claims 1, 11), "the graphene film" after "a graphene layer" (claims 8, 18), "the external rotor wheel" never introduced (roof drive claims 5-7), and the claim words "discrete" and "outer" that the specifications never use. 3 are false: "the structure" of the metal (inherent, twice) and "the heating step" after "selectively heating" (implicit).
- Support map on the same patents: 3 elements flagged that the labels call supported (all judgment calls, see 1).
4. Antecedent basis on random granted claims (code only)
Each round's flags were judged real or false (real = no explicit antecedent that a careful examiner could object to under MPEP 2173.05(e), arguable cases included; false = implicit or inherent antecedent, idiom, acronym, formula variable, superlative, spelling slip, or a phrase the extractor ran into a verb). The rules were then improved on that round, and the next round was drawn fresh; round 4 was drawn after the last change and is the held-out number. Judges: Claude for rounds 1-2, separate Claude agents given a written rubric for rounds 3-4.
| Round | Patents | Claims | Flags when judged | Genuine | Precision when judged | Precision with the final rules |
|---|---|---|---|---|---|---|
| 1 (dev) | 43 | 730 | 82 | 31 | 38% | 79% (30 of 31 genuine kept) |
| 2 (dev) | 55 | 920 | 113 | 52 | 46% | 81% (50 of 52) |
| 3 (dev) | 60 | 937 | 203 | 58 | 29% | 66% (56 of 58) |
| 4 (held out) | 55 | 891 | 66 | 36 | 55% | 55% |
- On unseen granted claims, about one flag in two is a genuine slip, at about 1.2 flags per patent. The rest are mostly references a practitioner would accept (a verb form, an inherent property, an idiom) and phrase-extraction errors. The check is lenient by design; it is a proofreading aid, and each flag shows the closest earlier term.
- Recall is measured only on planted errors (11 of 11 above); we did not label every reference in the random claims.
- Each round surfaced new false-alarm classes, so expect precision on other art units to differ.
Not measured
- Labels by a registered practitioner, or examiner reasons for allowance.
- Long specifications (over 48,000 characters) in retrieved mode; the eight eval patents all went in whole.
- The 32 GB card and the wanted GLM-5.3-Flash tier.
- Drafting help, reference numerals and figures (not built).