Skip to content
decosa

43 · Science and research · live

Citation and claim checker for papers

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Planted wrong-paper citations flagged12 of 12test splitn = 1295% CI 76-100%
  • Planted overstatements flagged12 of 12test splitn = 12All came back contradicted, with the differing passage quoted. 95% CI 76-100%.
  • False alarms on correct citations with open full text12 of 49 (24.5%)test splitn = 4995% CI 14.6-38.1%: about one flag in four is noise.
  • Planted retracted citations / year errors / DOI swaps flagged6 of 6 / 6 of 6 / 6 of 6test splitn = 18Year and DOI results are after fixes made following the test run (before: 5 of 6 and 4 of 6).
  • Retraction flag on known retracted works, with DOI / without68 of 68 / 68 of 68held outn = 68Drawn from Crossref's own notices: tests the pipeline, not Retraction Watch coverage.
  • Retraction flag on random control articles, with DOI / without0 of 60 / 0 of 60held outn = 60

Dataset

Eight CC BY PubMed Central papers: 2 dev papers used to write the prompt, 6 test papers fixed before any test run (471 citation pairs), with planted errors; 60 unplanted citation pairs labelled blind; a retraction set of 68 retracted works and 60 random controls from Crossref/OpenAlex metadata.

Caveats

  • False-alarm labels are an AI agent's, not a domain expert's, on only 49 correct citations: a wide interval.
  • Overstatement plants were hand-written by the builder.
  • Only open-access cited papers can be checked (44% of test pairs had full text); figures and tables are not read.
  • Reference-check fixes were made after the test run and re-run without the model; the claim verdicts are from the first run.
  • Run-to-run variation: the same sentence can be flagged in one run and supported in the next.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
9.6 s
Receipts
4
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume plus the GROBID container, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Verified on 2026-09-25: the image builds, the service starts healthy with GROBID parsing, the smoke test passes (10.5 s, 4 attested calls), the planted sample flags reference 43 retracted, the overstated [42] sentence and reference 1's year, the signed report verifies and fails when one verdict is changed, and no manuscript text reaches the logs. In one of three runs the swapped [36] citation came back supported.

Rehearsal bundle: paper-claim-check.zip (2 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • About one flag in four on correct background citations is noise (12 of 49 on the test papers); each flag quotes the passage, so a reader can dismiss it quickly.
  • Only open-access cited papers are read; 56% of test citations had only an abstract or nothing, and those are marked unavailable.
  • Labels and plants are Claude's (an AI agent), not a domain expert's, on six biomedical papers.
  • Run-to-run variation: the same citation can be flagged in one run and supported in the next.
  • PDF text loses superscript citation numbers; figures and tables are not read.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)decosa-api papercheck module (decosa_api/verticals/papercheck)AGPL-3.0-or-later
  • Reference parser (optional): splits each reference into authors, title, year and DOIGROBID 0.8.2 (CRF models)Apache-2.0
  • Model: reads each citing sentence against the cited paper's passages, then reviews its own flagsQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · references only, any CPU (4)
  • Retraction flag on known retracted works (60 sampled from Crossref's notices + 8 well-known), cited with DOI / without a DOI: 68 of 68 / 68 of 68decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the sampled set comes from the same Crossref data, so this tests the pipeline, not Retraction Watch's coverage
  • Retraction flag on 60 random journal articles (controls), with DOI / without: 0 of 60 / 0 of 60decosa-api docs/evals/paper-claim-check.md, 2026-09-25
  • Planted retracted citations / wrong years / wrong DOIs found in 6 test papers: 6 of 6 / 6 of 6 / 6 of 6decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the no-DOI retraction and the DOI plants after fixes made on the test run (disclosed there)
  • Unplanted reference flags on the same 6 real papers (431 references): genuine / arguable / wrong: 3 / 3 / 0 errors and warnings (two wrong DOIs in a published paper, one malformed author list; one ambiguous supplement DOI, two software records whose year differs), after the fixesdecosa-api docs/evals/paper-claim-check.md, 2026-09-25; judged by Claude (an AI agent)
Standard · one 96 GB card (measured; hosted demo) (3)
  • Planted wrong-paper citations flagged not supported or contradicted, 6 held-out CC BY papers: 12 of 12decosa-api docs/evals/paper-claim-check.md, 2026-09-25
  • Planted overstatements flagged (numbers inflated, association made causal, population or design changed): 12 of 12, each with the differing passage quoteddecosa-api docs/evals/paper-claim-check.md, 2026-09-25; plants written by Claude (an AI agent)
  • False alarms on correct citations with open full text (blind labels): 12 of 49 flagged (24.5%, 95% CI 14.6-38.1%)decosa-api docs/evals/paper-claim-check.md, 2026-09-25; labels written by Claude (an AI agent), not domain experts, before the verdicts were seen
Wanted · GLM-5.3-Flash for the claim check (1)
  • This eval, same protocol: not measured yet

How we measure · All tools