43 · Science and research · live
Citation and claim checker for papers
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Planted wrong-paper citations flagged12 of 12test splitn = 1295% CI 76-100%
- Planted overstatements flagged12 of 12test splitn = 12All came back contradicted, with the differing passage quoted. 95% CI 76-100%.
- False alarms on correct citations with open full text12 of 49 (24.5%)test splitn = 4995% CI 14.6-38.1%: about one flag in four is noise.
- Planted retracted citations / year errors / DOI swaps flagged6 of 6 / 6 of 6 / 6 of 6test splitn = 18Year and DOI results are after fixes made following the test run (before: 5 of 6 and 4 of 6).
- Retraction flag on known retracted works, with DOI / without68 of 68 / 68 of 68held outn = 68Drawn from Crossref's own notices: tests the pipeline, not Retraction Watch coverage.
- Retraction flag on random control articles, with DOI / without0 of 60 / 0 of 60held outn = 60
Dataset
Eight CC BY PubMed Central papers: 2 dev papers used to write the prompt, 6 test papers fixed before any test run (471 citation pairs), with planted errors; 60 unplanted citation pairs labelled blind; a retraction set of 68 retracted works and 60 random controls from Crossref/OpenAlex metadata.
Caveats
- False-alarm labels are an AI agent's, not a domain expert's, on only 49 correct citations: a wide interval.
- Overstatement plants were hand-written by the builder.
- Only open-access cited papers can be checked (44% of test pairs had full text); figures and tables are not read.
- Reference-check fixes were made after the test run and re-run without the model; the claim verdicts are from the first run.
- Run-to-run variation: the same sentence can be flagged in one run and supported in the next.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 9.6 s
- Receipts
- 4
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume plus the GROBID container, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.
Verified on 2026-09-25: the image builds, the service starts healthy with GROBID parsing, the smoke test passes (10.5 s, 4 attested calls), the planted sample flags reference 43 retracted, the overstated [42] sentence and reference 1's year, the signed report verifies and fails when one verdict is changed, and no manuscript text reaches the logs. In one of three runs the swapped [36] citation came back supported.
Rehearsal bundle: paper-claim-check.zip (2 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- About one flag in four on correct background citations is noise (12 of 49 on the test papers); each flag quotes the passage, so a reader can dismiss it quickly.
- Only open-access cited papers are read; 56% of test citations had only an abstract or nothing, and those are marked unavailable.
- Labels and plants are Claude's (an AI agent), not a domain expert's, on six biomedical papers.
- Run-to-run variation: the same citation can be flagged in one run and supported in the next.
- PDF text loses superscript citation numbers; figures and tables are not read.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)decosa-api papercheck module (decosa_api/verticals/papercheck)AGPL-3.0-or-later
- Reference parser (optional): splits each reference into authors, title, year and DOIGROBID 0.8.2 (CRF models)Apache-2.0
- Model: reads each citing sentence against the cited paper's passages, then reviews its own flagsQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · references only, any CPU (4)
- Retraction flag on known retracted works (60 sampled from Crossref's notices + 8 well-known), cited with DOI / without a DOI: 68 of 68 / 68 of 68decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the sampled set comes from the same Crossref data, so this tests the pipeline, not Retraction Watch's coverage
- Retraction flag on 60 random journal articles (controls), with DOI / without: 0 of 60 / 0 of 60decosa-api docs/evals/paper-claim-check.md, 2026-09-25
- Planted retracted citations / wrong years / wrong DOIs found in 6 test papers: 6 of 6 / 6 of 6 / 6 of 6decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the no-DOI retraction and the DOI plants after fixes made on the test run (disclosed there)
- Unplanted reference flags on the same 6 real papers (431 references): genuine / arguable / wrong: 3 / 3 / 0 errors and warnings (two wrong DOIs in a published paper, one malformed author list; one ambiguous supplement DOI, two software records whose year differs), after the fixesdecosa-api docs/evals/paper-claim-check.md, 2026-09-25; judged by Claude (an AI agent)
Standard · one 96 GB card (measured; hosted demo) (3)
- Planted wrong-paper citations flagged not supported or contradicted, 6 held-out CC BY papers: 12 of 12decosa-api docs/evals/paper-claim-check.md, 2026-09-25
- Planted overstatements flagged (numbers inflated, association made causal, population or design changed): 12 of 12, each with the differing passage quoteddecosa-api docs/evals/paper-claim-check.md, 2026-09-25; plants written by Claude (an AI agent)
- False alarms on correct citations with open full text (blind labels): 12 of 49 flagged (24.5%, 95% CI 14.6-38.1%)decosa-api docs/evals/paper-claim-check.md, 2026-09-25; labels written by Claude (an AI agent), not domain experts, before the verdicts were seen
Wanted · GLM-5.3-Flash for the claim check (1)
- This eval, same protocol: not measured yet