Skip to content
decosa

36 · Legal · live

Patent claim-support checker

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Supported-or-not call agrees with the labels93.4%test splitn = 166100% on the 146 elements labelled as clear calls. Three runs: 93.6%, 91.0%, 93.4%. Dev: 97.8%.
  • Supported elements where a cited paragraph is one the labels list100%test splitn = 150Share of cited paragraphs the labels list: 93.8%
  • Partly supported elements flagged (recall) / flags that match a label (precision)5 of 13 / 5 of 8test splitn = 166Partial support is what it misses; all misses and false flags are judgment calls.
  • Planted support removal flagged5 of 6test splitn = 6Over two runs 10 of 12; both misses cite a paragraph that still describes the element.
  • Antecedent-basis flags that are genuine, random granted claims55% (36 of 66)held outn = 66Round 4: 55 patents, 891 claims, drawn after the last rule change. About 1.2 flags per patent.
  • Cost per application (list price)$0.053test split166 calls over 6 patents, about 6,000 prompt tokens per element

Dataset

Eight granted US patents (public domain): 2 dev, 6 test (212 claim elements; 166 test), labelled per element by two separate Claude agents. Antecedent set: 213 randomly drawn granted patents in four rounds (3,478 claims); rounds 1-3 used to improve the rules, round 4 held out.

Caveats

  • Labels are by AI agents, not a registered practitioner; no examiner reasons for allowance to check against.
  • Granted claims were examined, so most elements are expected to be supported.
  • Partial support is often missed (5 of 13): treat "supported" as "here is where to look", not "no 112(a) problem".
  • Antecedent recall is measured only on planted errors; precision may differ in other art units.
  • Not measured: long specifications over 48,000 characters, the 32 GB card tier.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
7.8 s
Receipts
13
Model calls
n/a
Tokens
n/a
Cost per run
$0.020

Self-host verification

Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Verified on 2026-09-25: the image builds, the service starts healthy, the planted sample runs end to end on the direct route (13 attested calls, 3.4 s): claims 2 and 5 have no support, the antecedent, reference and term issues are found, the signed record verifies and fails when one strength is changed, the CSV export works, and no claim text reaches the logs. The model server's own startup was not re-verified (no new GPU load).

Rehearsal bundle: patent-claim-support.zip (15 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Labels are Claude's (an AI agent), not a registered practitioner's; antecedent flags were judged by separate Claude agents with a written rubric.
  • The support map finds supporting paragraphs well but flags only 5 of the 13 held-out elements the labels call partly supported; most misses are judgment calls such as a feature described only in another embodiment.
  • About half of the antecedent-basis flags on randomly drawn granted claims are genuine slips (36 of 66 held out); the rest are implicit or inherent references a practitioner would accept.
  • Run-to-run variation on the hosted gateway: the same element can come back partly supported in one run and unsupported in the next; both are flagged.
  • Text only (no DOCX, PDF or figures); specifications over 48,000 characters are read as the best-matching paragraphs, which can miss support far from the element's words.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)decosa-api patent module (decosa_api/verticals/patent)AGPL-3.0-or-later
  • Model: reads each claim element against the specification and names the supporting paragraphsQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · code checks only, any CPU (2)
  • Antecedent-basis flags that are genuine, on randomly drawn granted US claims (held out: rules frozen before the set was drawn): 54.5% (36 of 66 flags; 55 patents, 891 claims)decosa-api docs/evals/patent-claim-support.md, 2026-09-25; each flag judged by a separate Claude agent with a written rubric, not a practitioner
  • Planted defects found by code in 6 test patents: antecedent errors / broken claim references / renamed terms: 11 of 11 / 12 of 12 / 6 of 6decosa-api docs/evals/patent-claim-support.md, 2026-09-25
Standard · one 96 GB card (measured; hosted demo) (5)
  • Elements where the model's supported-or-not call matches the labels, 6 held-out granted patents: 93.4% of 166 (100.0% of the 146 labelled as clear calls)decosa-api docs/evals/patent-claim-support.md, 2026-09-25; labels written by Claude (an AI agent) reading each specification, not by a practitioner
  • Supported elements where a cited paragraph is one the labels list / share of cited paragraphs that are: 100.0% / 93.8% (150 elements)decosa-api docs/evals/patent-claim-support.md, 2026-09-25
  • Elements the labels call only partly supported that the model also flags (recall) / flags that match a label (precision): 5 of 13 / 5 of 8decosa-api docs/evals/patent-claim-support.md, 2026-09-25; most misses are limitations the labels mark as judgment calls, such as a feature described only in a different embodiment
  • Planted support removal: all paragraphs describing one element deleted, element flagged: 5 of 6 (the miss: a paragraph the removal left in still describes the element, checked by reading it; an earlier run missed a different patent the same way)decosa-api docs/evals/patent-claim-support.md, 2026-09-25
  • Dev patents (2, used to write the prompt): agreement / flag recall: 97.8% / 3 of 3decosa-api docs/evals/patent-claim-support.md, 2026-09-25
Wanted · a GLM-5.3-Flash judge on your own hardware (1)
  • This eval, same protocol: not measured yet

How we measure · All tools