Eval: citation and claim checker for papers (43)
Run on our server, 25 Sep 2026. Model: Qwen3.8-27B (NVFP4) through the model gateway, temperature 0, thinking off, every call
receipted, at most 2 calls in flight (the gateway was shared and saturated). Data and scripts:
docs/evals/paper-claim-check/, scripts/papercheck_eval.py, scripts/papercheck_data.py.
Data and licences
- Eight open-access papers from PubMed Central (Europe PMC JATS), all CC BY: two dev papers used to write the prompt
(PMC8389398, PMC7326225) and six test papers fixed before any test run (PMC8675677, PMC7732057, PMC8321109,
PMC7909686, PMC8196915, PMC7055896; PLoS Comput Biol, PLoS Med, BMC Public Health, PLoS One; COVID-19, tuberculosis,
malaria, vaccines, sleep and vitamin D). Each is rendered as title, abstract, introduction, discussion and the full
reference list, with every in-text citation as
[n]. The licence of each is checked inbuild. - Cited papers: whatever Europe PMC holds as open access (full text) or as an abstract; arXiv for preprints. Of the 471 citation pairs in the six test papers, 209 had open full text, 224 only an abstract and 38 nothing.
- Retraction set: 60 retracted works sampled at random from Crossref's retraction notices, 8 well-known retractions picked from memory (Wakefield 1998, the two Surgisphere papers of 2020, STAP 2014, Hwang 2005, LaCour 2014, Séralini 2012, Macchiarini 2011) and 60 random journal articles (2005–2022) as controls. Metadata is Crossref's and OpenAlex's (CC0); the Retraction Watch data is openly available through Crossref.
Politeness and rate limits used
User-Agent decosa-paper-claim-check/1.0 (https://decosa.ai; mailto:<contact>) and mailto= on every
Crossref and OpenAlex call. Per-host limits, shared by the whole process: Crossref works 5/s, 3 concurrent (Crossref's
own headers on 25 Sep 2026: polite-single 10/s, 3 concurrent); Crossref bibliographic queries 2/s, 2 concurrent
(polite-array allows 3/s); OpenAlex 5/s, singleton lookups by DOI only (free under OpenAlex's pricing; searches and lists
cost credits, so none are made); Europe PMC 5/s; arXiv one request every 3 s; doi.org agency lookups 2/s, only for DOIs
Crossref does not know. 429 and 5xx are retried twice, honouring Retry-After. Public metadata and open-access text
were cached on disk for the eval. The committed runs and labelling sheet keep verdicts, ids and titles but not the cited
papers' passages or abstracts (not all cited papers are CC BY), so repeat runs made no new requests; the retraction eval made 392 requests the first
time (Crossref 263, OpenAlex 129).
Protocol
- Prompt written on the two dev papers only (with four dev plants), including a second "review" call that keeps a PARTIAL or CONTRADICTED flag only when the model quotes a passage that is really in the cited paper's text.
- Plants (
plan.json), per test paper: two wrong-paper swaps (a single-citation marker re-pointed at the open-full-text reference least like the sentence, seeded), two hand-written overstatements (numbers inflated, association turned into cause, population or design changed; written from the sentence and the cited title only), one citation to a retracted paper (half with DOI, half without), a year three years off and a DOI replaced by another reference's DOI. - False-alarm labels (
labels.json): 60 unplanted citation pairs with open full text (10 per test paper, seeded), labelled by Claude (an AI agent, not a domain expert) from a sheet with the sentence, the cited title, abstract and best-matching passages, before the checker's test verdicts were seen. 49 were labelled correct citations, 7 unsure (excluded) and 4 were planted sentences that leaked into the sheet (excluded). - The checker ran over HTTP on the planted papers (real receipts), then
score.
Results
| Measure | Result |
|---|---|
| Wrong-paper swaps flagged unsupported or contradicted | 12 of 12 (95% CI 76–100%) |
| Overstatements flagged (all came back contradicted, with the differing passage quoted) | 12 of 12 (76–100%) |
| Planted retracted citations flagged | 6 of 6 (3 cited without a DOI) |
| Planted year errors / DOI swaps flagged | 6 of 6 / 6 of 6 (after fixes 3 and 4 below; before them 5 of 6 and 4 of 6) |
| False alarms on correct citations with open full text | 12 of 49 flagged (24.5%, 95% CI 14.6–38.1%): 8 unsupported, 3 contradicted, 1 partial |
| Retraction flag, known retracted works (random 60 + famous 8), cited with DOI / without | 68 of 68 / 68 of 68 |
| Retraction flag on controls (60 random articles), with DOI / without | 0 of 60 / 0 of 60 (57 of 60 matched without a DOI; 3 unmatched) |
OpenAlex is_retracted agrees with the Crossref-based flag |
on every matched work, retracted and control |
What the false alarms are. Mostly background citations the model reads too literally ("the cited paper does not state the general definition ..."), clause attribution in long sentences with several citations, and one where the cited paper is a trial cited as "literature on diagnostic performance". Three of the twelve are arguable (for example a median of 55 nmol/L cited for "insufficiency below 50 nmol/L is common"). So roughly one in four flags on a real paper is noise that the quoted passage lets a reader dismiss in seconds; the plants are caught every time. The review call removed 23 first-pass flags across the test papers.
Retraction flags are not an independent test of coverage. The retracted set was drawn from Crossref's own notices, so this measures the pipeline (reference string → record → notice), not how complete Retraction Watch is. The hard case is a reference without a DOI whose title also exists as a preprint, a duplicate record or a notice: Wakefield 1998 has an un-retracted duplicate DOI in Crossref, and the checker also looks for retracted records with the same title.
Unplanted reference flags on the six real papers (431 references). Errors and warnings: two wrong DOIs in a published paper (PMC8675677 refs 43 and 45 point to "Big hopes for big data" and a cancer-cell paper; the checker names the right DOIs), one malformed author list (given names printed as surnames), one ambiguous DOI (a journal supplement whose DOI record is titled "1. General introduction"), and two year warnings on software records (R, a CRAN package) whose Crossref year differs. Notes: six one-year differences (online vs print), two DOIs registered with DataCite or mEDRA, and self-citation stacks of 3 and 4 papers in two papers.
Fixes made after the test run, disclosed. (1) A DOI Crossref does not know was reported as "not registered"; two
real references use DataCite and mEDRA DOIs, so the checker now asks doi.org for the registration agency and reports
those as a note. (2) A search match accepted when the reference had no author (a CDC web page matched a 2024 record);
now a match needs the year (±1) or a real author match. (3) Retraction prefixes ("RETRACTED ARTICLE:") and preprint
twins lowered the title match, so the no-DOI retracted plant in PMC7055896 matched the preprint; titles are compared
without notice prefixes and a journal article with the exact year wins. (4) Two planting bugs in the eval script (DOI
case; a marker added by another plant hid a sentence). The claim judge and its prompt were not changed after the test
run; the reference checks were re-run without the model (refs), the claim verdicts are from the first run.
Cost and latency (measured, shared gateway, 2 calls in flight)
- Test papers: 29–143 citations checked each; 504 model calls, 855,890 prompt and 37,690 generated tokens, $0.313 for the six papers ($0.021–0.098 each) at the gateway's list price ($0.30 / $1.50 per million tokens). 35–247 s per paper.
- Demo sample (24 citations, 43 references): 19.8 s, 25 calls, $0.017; the published version 13.7 s, 22 calls, $0.015.
- Run-to-run variation: the same sentence can be flagged in one run and supported in the next (dev paper, "a whole country" cited to a city-scale study: flagged in four runs, supported in one).
Limits
Labels are an AI agent's, on 49 correct citations: a small set with a wide interval. Only open-access cited papers can be checked (44% of test pairs had full text); abstracts settle support and contradiction only, never absence. Figures and tables are not read. The model can misattribute a clause in a long multi-citation sentence.
Verdict (value check)
An editorial office or integrity team would pay for this as a screening pass: the reference checks (wrong DOIs, retractions, notices, self-citation stacks) are precise and cheap, and the claim check catches wrong-paper citations and changed findings reliably with the passage quoted. What is missing for a paid pilot: fewer background false alarms (about one flag in four is noise), labels from domain experts, access to subscription full text through the publisher's own licence (TDM) on the customer's box, and a Word or manuscript-system (ScholarOne, Editorial Manager) integration.