Check a paper's citations
An issue list where each citing sentence is checked against the cited paper's open full text, with the passage quoted, plus retracted or wrong sources.
Built on: Grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the citation and claim checker for papers API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the reference checks run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- paper-claim-check
Use the hosted API
# Decosa citation and claim checker: use the hosted API
You are wiring Decosa's citation and claim checker into this project. It takes a manuscript with its reference list and
returns: every reference checked against Crossref and OpenAlex (retracted, withdrawn, corrected, a DOI that does not
exist or points to another paper, a wrong year or first author), every citing sentence checked against the cited
paper's open full text with the passage quoted, self-citation stacks, and a signed report. Each model call has its own
signed receipt. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for published papers, your own drafts and testing your integration. A manuscript under peer review
is confidential: for that, use the self-host prompt instead.
- It is a screening aid, not a finding of misconduct. Show every flag as something for a person to check.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "paper-claim-check"}` returns `{"token", "expires_at", "budget"}`.
a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about 110 checked citations). Over a limit: HTTP 429 with `Retry-After`.
3. One check at a time per demo token (409 otherwise).
## Endpoints
- `POST /papercheck/extract` (token) `{"filename": "paper.pdf", "file_b64": "..."}` → `{text, kind, references, citations, style, notes}`.
PDF, DOCX, LaTeX, text; up to about 11 MB. PDF text loses superscript citation numbers; prefer DOCX or LaTeX.
- `POST /papercheck/check` (token). Body: `{"text": "...", "authors"?: ["Jane Smith", "Kim Lee"], "title"?: "...", "check_claims"?: true, "stream"?: true}`
or `{"sample": "telecommuting-planted"}`.
- The reference list goes at the end under a heading such as "References". Citations: `[3]`, `[3-5]`, `(3)` or
author-year `(Smith et al., 2020)`.
- Limits: 250,000 characters, 250 references, 60 checked citations per request (the rest come back `skipped`).
`"check_claims": false` runs the reference checks only (no model call, no budget).
- With `"stream": true` it streams `ready` (the text, citations and references with offsets), a `reference` event per
reference, a `receipt` after each model call, a `claim` event per citation, then `report`, `done` and `budget`.
With `"stream": false`: one JSON object `{run_id, text, citations, receipts, report, budget}`.
- Claim verdicts: `supported`, `partial` (overstated or changed), `unsupported`, `contradicted`, `unavailable` (no open
full text, or only an abstract that cannot settle it), `error`, `skipped`. Each has `evidence` (quoted passages),
`reason`, `basis` and `receipt_ids`.
- `report.issues`: `[{id, kind, severity: error|warning|note, message, ref?, claim?}]`; kinds include `retracted`,
`expression_of_concern`, `corrected`, `doi_not_found`, `doi_mismatch` (with `suggested_doi`), `year_mismatch`,
`author_mismatch`, `missing_reference`, `not_cited`, `self_citation_stack`, `claim`.
- `GET /papercheck/runs/{run_id}/export?format=md|json|record`.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.
- `GET /papercheck/info`, `GET /papercheck/samples`, `GET /papercheck/samples/{id}`, `GET /attest/signing-key` (no token).
## Example: check a paper and print the flags (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
text = open("paper.txt").read() # or POST /papercheck/extract for a PDF or DOCX
r = httpx.post(f"{API}/papercheck/check", headers=H, timeout=900,
json={"text": text, "authors": ["Jane Smith"], "stream": False})
r.raise_for_status()
rep = r.json()["report"]
for x in rep["issues"]:
if x["severity"] != "note":
print(x["severity"], x["kind"], x["message"])
```
## Honest limits
- Only open-access cited papers are read; in our test papers 56% of citations had only an abstract or nothing, and
those come back `unavailable`.
- About one flag in four on correct background citations is noise (measured 12 of 49); each flag quotes the passage.
- Retraction data comes from Crossref (with the Retraction Watch data) and OpenAlex and can lag the publishers.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa citation and claim checker: run it yourself (containers)
You are setting up the Decosa citation and claim checker on this machine, so confidential manuscripts never leave it.
It checks every reference against Crossref and OpenAlex (retractions, corrections, wrong DOIs and years), every citing
sentence against the cited paper's open full text, flags self-citation stacks and signs a report. Only DOIs and
reference strings go to the public metadata services; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/paper-claim-check.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py paper-claim-check` (the api image carries the same bundle under /app/rehearsal/paper-claim-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py paper-claim-check --bundle paper-claim-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "four references are found", "citation [1] is supported with a quoted passage", "citation [2] is supported with a quoted passage"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; add `lfoppiano/grobid:0.8.2` as a
`grobid` service if I want better reference parsing. For the `api` service set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_PAPERCHECK_MAILTO=<my contact address>`
(ask me for it), `DECOSA_PAPERCHECK_GROBID_URL=http://grobid:8070` if GROBID is there, and bind every port to
127.0.0.1. The api needs outbound HTTPS to api.crossref.org, api.openalex.org, www.ebi.ac.uk, arxiv.org and doi.org.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/papercheck/info` shows `"route": "direct"` and my contact address in the
User-Agent; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify
my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"paper-claim-check"}` and run
`POST /papercheck/check {"sample": "telecommuting-planted", "stream": false}`. Expect reference 43 flagged retracted,
the `[36]` sentence `unsupported`, the `[42]` sentence `contradicted` with a quoted passage and a `year_mismatch` on
reference 1; every receipt `attested`. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the flags and how long the run took.
Off, and keep it off on a box that holds confidential manuscripts: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? Parts of this tool also run on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/paper-claim-check-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMlite tierRuns with a smaller tier
The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested)
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested)
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"paper-claim-check"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py paper-claim-check
Download the mock-data bundle (2 KB, 10 checks)expected.json
A four-sentence manuscript citing three papers from the reference list of an open-access CC BY article, plus the retracted 2020 hydroxychloroquine registry study. Citations [1] and [2] must be supported with a quoted passage, [3] (a hospital Staphylococcus aureus paper cited for SARS-CoV-2 super-spreaders) must be flagged, reference 4 must be flagged retracted, and the signed record must verify and catch a changed entry. Lookups go to Crossref, OpenAlex and Europe PMC.
What the rehearsal checks
- four references are found
- citation [1] is supported with a quoted passage
- citation [2] is supported with a quoted passage
- the wrong paper cited as [3] is flagged
- reference 4 is flagged retracted
- the retraction is raised as an error issue
- the signed record verifies
- a record with the [3] verdict changed to supported no longer verifies
- verification points at a claim entry
- every model call has a signed receipt
Licence: The sentences are written for Decosa; the three reference strings come from the reference list of Mauras S et al., PLoS Comput Biol 2021;17(8):e1009264 (CC BY 4.0), and reference 4 is the public bibliographic entry of Mehra et al., The Lancet 2020 (retracted). Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa citation and claim checker: run it yourself (containers)
You are setting up the Decosa citation and claim checker on this machine, so confidential manuscripts never leave it.
It checks every reference against Crossref and OpenAlex (retractions, corrections, wrong DOIs and years), every citing
sentence against the cited paper's open full text, flags self-citation stacks and signs a report. Only DOIs and
reference strings go to the public metadata services; nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/paper-claim-check.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py paper-claim-check` (the api image carries the same bundle under /app/rehearsal/paper-claim-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py paper-claim-check --bundle paper-claim-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "four references are found", "citation [1] is supported with a quoted passage", "citation [2] is supported with a quoted passage"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; add `lfoppiano/grobid:0.8.2` as a
`grobid` service if I want better reference parsing. For the `api` service set `DECOSA_LLM_ROUTE=direct`,
`DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_PAPERCHECK_MAILTO=<my contact address>`
(ask me for it), `DECOSA_PAPERCHECK_GROBID_URL=http://grobid:8070` if GROBID is there, and bind every port to
127.0.0.1. The api needs outbound HTTPS to api.crossref.org, api.openalex.org, www.ebi.ac.uk, arxiv.org and doi.org.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/papercheck/info` shows `"route": "direct"` and my contact address in the
User-Agent; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify
my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"paper-claim-check"}` and run
`POST /papercheck/check {"sample": "telecommuting-planted", "stream": false}`. Expect reference 43 flagged retracted,
the `[36]` sentence `unsupported`, the `[42]` sentence `contradicted` with a quoted passage and a `year_mismatch` on
reference 1; every receipt `attested`. Then `POST /record/verify` with `report.record`: `ok` must be true.
6. Report back: the public key and key id, the totals, the flags and how long the run took.
Off, and keep it off on a box that holds confidential manuscripts: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? Parts of this tool also run on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/paper-claim-check-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsCitation and claim checker for papers on GeForce RTX 5090: use the Standard · one 96 GB card (measured; hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts of about 3,000 tokens. Estimate: same model as the measured card, not run here on a 5090.
Standard · one 96 GB card (measured; hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Checker: decosa-api papercheck module (decosa_api/verticals/papercheck). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Reference parser: GROBID 0.8.2 (CRF models). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Citation and claim checker for papers, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Citation and claim checker for papers on my hardware Fetch https://decosa.ai/prompts/paper-claim-check-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=paper-claim-check) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one 96 GB card (measured; hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Checker: decosa-api papercheck module (decosa_api/verticals/papercheck), CPU - Reference parser: GROBID 0.8.2 (CRF models), CPU - Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/paper-claim-check-assemble.md
Partly on a Mac
Some parts run natively on Apple Silicon (32 GB or more); the rest needs a CUDA GPU or a hosted API. Measured speeds and what runs where
- GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested)
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Citation and claim checker for papers: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Citation and claim checker for papers on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Only part of this tool runs on a Mac (see the gaps below). The parts that do need 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/paper-claim-check.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py paper-claim-check` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "four references are found", "citation [1] is supported with a quoted passage", "citation [2] is supported with a quoted passage"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Reference parser (optional): splits each reference into authors, title, year and DOI | Java service on CPU (Docker) | Docker Desktop, amd64 image under emulation | Untested on a Mac | | Model: reads each citing sentence against the cited paper's passages, then reviews its own flags | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | Not on a Mac: - GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested). ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. The script sets up only the language model. The parts listed under "Not on a Mac" still need the GPU stack or the hosted API (see https://decosa.ai/prompts/paper-claim-check-assemble.md); the smoke test in step 6 will report them as failures. Tell me which checks passed and which need the GPU. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py paper-claim-check`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 9.6 s · ~$0.003 per run · 4 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the api service with a named volume plus the GROBID container, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.
Measured cost to run: about $0.066 per 100 citations checked (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Verified on 2026-09-25: the image builds, the service starts healthy with GROBID parsing, the smoke test passes (10.5 s, 4 attested calls), the planted sample flags reference 43 retracted, the overstated [42] sentence and reference 1's year, the signed report verifies and fails when one verdict is changed, and no manuscript text reaches the logs. In one of three runs the swapped [36] citation came back supported.
Known limits (5)
- About one flag in four on correct background citations is noise (12 of 49 on the test papers); each flag quotes the passage, so a reader can dismiss it quickly.
- Only open-access cited papers are read; 56% of test citations had only an abstract or nothing, and those are marked unavailable.
- Labels and plants are Claude's (an AI agent), not a domain expert's, on six biomedical papers.
- Run-to-run variation: the same citation can be flagged in one run and supported in the next.
- PDF text loses superscript citation numbers; figures and tables are not read.
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the citation and claim checker for papers API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the reference checks run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Every citing sentence checked against the cited paper's open full text, with the passage quoted; retracted and corrected sources, wrong DOIs and years, and self-citation stacks flagged; one receipt per check and a signed report.
Paste a manuscript with its reference list, or upload a PDF, DOCX or LaTeX file. Each reference is looked up in Crossref and OpenAlex: retractions, withdrawals, expressions of concern and corrections (Crossref's notices, which include the Retraction Watch data), DOIs that do not exist or point to another paper, and wrong years or first authors. For every sentence that cites something, an open model reads the cited paper's open full text (PubMed Central open access through Europe PMC, or arXiv) and says whether the paper supports it, overstates it, does not say it or says otherwise, quoting the passage. Citations whose full text is not open are marked unavailable, never guessed. With the authors' names it also flags self-citation stacks. The result is an issue list and a signed report with hashes, verdicts, cited DOIs and receipt ids but no manuscript text. For authors, reviewers, editors and research-integrity offices; it screens, a person decides.
- Deployment
- Hosted or self-host
- Regulatory
- A screening aid, not a finding of research misconduct. Citation manipulation, including self-citation stacking, is treated by journals under COPE's guidance, and an allegation is for the journal or institution to investigate. A manuscript under peer review is confidential: NIH prohibits its peer reviewers from using generative AI tools to analyse or critique applications (notice NOT-OD-23-149, June 2023), and many publishers tell reviewers not to upload manuscripts to AI tools, so reviewers and editors should run the self-hosted version on their own machine. Only DOIs, identifiers and reference strings are sent to Crossref, OpenAlex, Europe PMC, arXiv and doi.org; the manuscript's text is not. Open-access full text is read under each article's own licence and only short passages are quoted back. Retraction and correction data come from Crossref (which has made the Retraction Watch database openly available since September 2023) and OpenAlex, and can lag the publishers. Model licence: Apache-2.0 (Qwen3.8-27B). Checked 25 Sep 2026.
Text description
A manuscript (text, PDF, DOCX or LaTeX) goes to the checker, which runs on your own machine. Code splits it into sentences, in-text citations and the reference list (GROBID optional). Only DOIs and reference strings go to Crossref, OpenAlex, Europe PMC, arXiv and doi.org, which return records, retraction and correction notices, and the cited papers' open text. Qwen3.8-27B (Apache-2.0) reads each citing sentence against the best-matching passages of the cited paper through the grounding checker's spans and verdict parser, one call per citation plus a review call for flags. Outputs: an issue list (retractions, wrong DOIs and years, self-citation stacks), each citation with a verdict and the quoted passage, and a signed report with hashes, verdicts, cited DOIs and receipt ids but no manuscript text. On the hosted route each model call gets a receipt that our gateway countersigns.
At a glance
- Data retention
- The manuscript is held in memory for the request; the run (verdicts, issues, the text for exports) for one hour, for the token or key that made it. Nothing is written to disk; logs carry counts only.
- What leaves the box
- Self-hosted: only DOIs, identifiers and reference strings, to Crossref, OpenAlex, Europe PMC, arXiv and doi.org. Hosted demo: also the text, to the model through our gateway, so use it for published papers or your own drafts.
- Model cost per paper
- A few cents to about ten cents per paper at the gateway's list price, rising with the number of checked citations (measured); the metadata services are free.
- Input
- Text with a reference list under a heading such as "References", or a PDF, DOCX or LaTeX file; numbered [3] or author-year (Smith et al., 2020) citations. Hosted limits: 250,000 characters, 250 references, 60 checked citations per request.
- Output
- An issue list, every citation with a verdict and the quoted passage, a Markdown or JSON report, and a signed record verifiable at /record/verify.
- Politeness
- Identifies itself with a contact address to every service, stays under Crossref's published polite-pool limits, uses only OpenAlex lookups that cost no credits, and fetches arXiv at most once every 3 seconds.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Lite
references only, any CPU
Retractions, corrections, wrong DOIs, years and first authors, missing and uncited references, and self-citation stacks, with no model and no GPU (check_claims: false). No claim check.
- Models
- decosa-api papercheck module (decosa_api/verticals/papercheck)
- GROBID 0.8.2 (CRF models)
- Hardware
- Any CPU
- Quality evidence
- Retraction flag on known retracted works (60 sampled from Crossref's notices + 8 well-known), cited with DOI / without a DOI68 of 68 / 68 of 68decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the sampled set comes from the same Crossref data, so this tests the pipeline, not Retraction Watch's coverage
- Retraction flag on 60 random journal articles (controls), with DOI / without0 of 60 / 0 of 60decosa-api docs/evals/paper-claim-check.md, 2026-09-25
- Planted retracted citations / wrong years / wrong DOIs found in 6 test papers6 of 6 / 6 of 6 / 6 of 6decosa-api docs/evals/paper-claim-check.md, 2026-09-25; the no-DOI retraction and the DOI plants after fixes made on the test run (disclosed there)
- Unplanted reference flags on the same 6 real papers (431 references): genuine / arguable / wrong3 / 3 / 0 errors and warnings (two wrong DOIs in a published paper, one malformed author list; one ambiguous supplement DOI, two software records whose year differs), after the fixesdecosa-api docs/evals/paper-claim-check.md, 2026-09-25; judged by Claude (an AI agent)
- Latency
- measured: seconds per paper with a warm cache; first lookups are paced by the polite rate limits.
- Verification
- Proof: partialNo model calls, so no receipts; the report is signed by the instance.
- In the hosted demo
Standard
one 96 GB card (measured; hosted demo)
Qwen3.8-27B on an RTX PRO 6000 reads every citing sentence against the cited paper's open text and reviews its own flags; the reference checks run alongside. Also fits a 32 GB card (estimate).
- Models
- decosa-api papercheck module (decosa_api/verticals/papercheck)
- GROBID 0.8.2 (CRF models)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB
- Quality evidence
- Planted wrong-paper citations flagged not supported or contradicted, 6 held-out CC BY papers12 of 12decosa-api docs/evals/paper-claim-check.md, 2026-09-25
- Planted overstatements flagged (numbers inflated, association made causal, population or design changed)12 of 12, each with the differing passage quoteddecosa-api docs/evals/paper-claim-check.md, 2026-09-25; plants written by Claude (an AI agent)
- False alarms on correct citations with open full text (blind labels)12 of 49 flagged (24.5%, 95% CI 14.6-38.1%)decosa-api docs/evals/paper-claim-check.md, 2026-09-25; labels written by Claude (an AI agent), not domain experts, before the verdicts were seen
- Latency
- measured on our server, shared gateway: under a minute to a few minutes per whole paper, rising with its citations; about half a minute for the demo.
- Verification
- Proof: strongHosted: every check and every review has its own gateway-signed receipt, listed in the signed report.
- Needs more compute
Wanted: the best setup
GLM-5.3-Flash for the claim check
A stronger reasoning model from another family for the background citations the standard model flags too eagerly. It needs two 96 GB cards or a large Mac, so community providers are welcome. Not served yet.
- Models
- decosa-api papercheck module (decosa_api/verticals/papercheck)
- GROBID 0.8.2 (CRF models)
- Qwen3.8-27B (NVFP4)
- GLM-5.3-Flash
- Hardware
- Network providers: 2x 96 GB cards or a Mac with 192 GB or more (about 170 GB of weights). Estimate.
- Quality evidence
- This eval, same protocolnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyNot hosted yet, so no receipts today.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)decosa-api papercheck module (decosa_api/verticals/papercheck) 0 GBProof: partial | LiteStandardWanted | 0 GB | Proof: partial | |
| ||||
Reference parser (optional): splits each reference into authors, title, year and DOIGROBID 0.8.2 (CRF models) 0 GBNo proof yet | LiteStandardWanted | 0 GB | No proof yet | |
| ||||
Model: reads each citing sentence against the cited paper's passages, then reviews its own flagsQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | StandardWanted | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Stronger open model for the claim check (wanted)GLM-5.3-Flash about 170 GB (estimate)No proof yetSelf-host only | Wanted | about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- Crossref REST API (opens in a new tab)Crossref metadata is open (no restrictions on reuse of the bibliographic facts); retraction and correction notices include the Retraction Watch data
Works by DOI (title, authors, year, updated-by notices) and bibliographic search for references without a DOI. Polite pool: User-Agent and mailto; our limits 5 requests/s for works and 2/s for searches, under Crossref's 10/s and 3/s.
Single works by DOI only (no credit cost): is_retracted as a second opinion, PMCID and open-access locations, the abstract. 5 requests/s.
- Europe PMC REST API (opens in a new tab)Open-access full text under each article's own licence (the demo paper is CC BY 4.0)
PMCID, open-access flag, abstract, and JATS full text of open-access articles. 5 requests/s.
- arXiv (opens in a new tab)Per paper, as chosen by the authors
PDF full text of cited preprints, at most one request every 3 seconds.
- scripts/papercheck_eval.py and docs/evals/paper-claim-check.mdApache-2.0
Builds the eval from six CC BY PMC papers, plants miscitations, scores detection, false alarms against blind labels, and retraction flags on 128 known retracted works and 60 controls.
- POST /record/verifyApache-2.0
Checks the signed report and names the first entry that was changed. The console also verifies it in your browser.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /papercheck/info, /papercheck/samples; POST /papercheck/extract (PDF, DOCX, LaTeX, text), /papercheck/check (SSE or JSON); GET /papercheck/runs/{id}/export?format=md|json|record. Keeps the manuscript in memory for the request and the run for one hour, never on disk; logs counts only.
- vLLM:8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
- GROBID (optional):8070
lfoppiano/grobid:0.8.2Reference parsing, CPU only; set DECOSA_PAPERCHECK_GROBID_URL.
Hardware
- Any CPU (references only) Fits
Measured: the reference checks for a 98-reference paper take 1-2 s with a warm cache; the first run is paced by the polite rate limits (estimate: well under a second per reference). No GPU.
- 1x RTX 5090 32 GB Fits
Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for prompts of about 3,000 tokens. Estimate: same model as the measured card, not run here on a 5090.
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the eval, the hosted demo and the self-host check ran on this card, shared with other services the whole time.
Latency per lane
- Demo sample, 24 citations and 43 references, hosted gateway route (2 calls in flight)19.8 s
Measuredmeasured on our server 2026-09-25, one run; the published version (23 citations) took 13.7 s
- Smoke manuscript, 4 citations, hosted gateway route9.6 s
Measuredmeasured on our server 2026-09-25: 8.8-10.4 s over 3 runs
- Whole test papers, 29-143 citations, hosted gateway route (2 calls in flight, saturated gateway)110.0 s
Measuredmeasured on our server 2026-09-25: 35-247 s over 6 papers
- Demo sample, self-hosted (clean clone, container build, direct route, GROBID on)32.6 s
Measuredmeasured on our server 2026-09-25, one run, 2 calls in flight on a shared card
Notes
- Only open-access cited papers can be read: in the six test papers 44% of citations had open full text, 48% only an abstract and 8% nothing. An abstract can confirm or contradict a citation but not show that the paper never says something, so those come back unavailable.
- A flagged verdict (overstated or contradicted) gets a second model call that must quote the passage that differs; a flag whose quote is not really in the paper is cleared. This removed 23 first-pass flags on the test papers without losing a planted one.
- Each reference without a DOI is matched by Crossref's bibliographic search and accepted only when most of the record's title is in the reference and the year or first author matches; a preprint twin or a duplicate record is handled, and a retracted record with the same title is flagged too.
- PDF text loses superscript citation numbers; paste text with [n] markers, or upload the DOCX or LaTeX.
- Not included yet: paywalled full text (a journal can run it with its own text-and-data-mining access on its own box), figures and tables, and a manuscript-system plug-in.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa citation and claim checker on this machine
You are setting up a citation checker for scientific manuscripts. For each reference it looks up the public record
(Crossref, OpenAlex): retractions, withdrawals, expressions of concern and corrections, DOIs that do not exist or point
to another paper, wrong years or first authors. For each sentence that cites something, an open model reads the cited
paper's open full text (Europe PMC open access, arXiv) and says whether it supports the sentence, with the passage
quoted. It also flags self-citation stacks and signs a report with this box's key. Work step by step, show me each
command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/paper-claim-check.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py paper-claim-check` (the api image carries the same bundle under /app/rehearsal/paper-claim-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py paper-claim-check --bundle paper-claim-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "four references are found", "citation [1] is supported with a quoted passage", "citation [2] is supported with a quoted passage"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Models and code: Qwen3.8-27B (Apache-2.0) as the judge; decosa-api (AGPL-3.0-or-later); optional GROBID 0.8.2 (Apache-2.0,
CPU only) for reference parsing.
- The manuscript stays on this machine. Only DOIs, PMIDs, arXiv ids and reference strings go out, to Crossref, OpenAlex,
Europe PMC, arXiv and doi.org. The service writes nothing about the manuscript to disk and logs counts only. Keep it
that way; do not add request logging. Bind every port to 127.0.0.1.
- Be polite to the public services: set `DECOSA_PAPERCHECK_MAILTO` to a contact address I give you (it goes in the
User-Agent and the `mailto` parameter). The built-in limits stay under Crossref's polite pool (5 requests/s for works,
2/s for searches), use only OpenAlex lookups by DOI (no credit cost) and fetch arXiv at most once every 3 seconds.
Do not raise them.
- Be honest about what it does: "supported" means the cited paper's open text backs the sentence as judged by a model,
not that the claim is true; "unavailable" means the full text is not open, not that the citation is wrong. It is a
screening aid; a person decides.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
6000 96 GB or an RTX 5090 32 GB works). Driver 570 or newer. Blackwell cards run NVFP4; on older cards use the FP8
weights.
2. No GPU? The reference checks (retractions, DOIs, years, self-citation) still run on CPU with `check_claims: false`;
there is no claim check and no model receipt.
3. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
4. Outbound HTTPS to api.crossref.org, api.openalex.org, www.ebi.ac.uk, arxiv.org and doi.org. Disk: about 30 GB.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
`decosa_api/verticals/papercheck/`, and `docker build -f docker/api/Dockerfile -t decosa-api:local .` (it includes
poppler's pdftotext for PDF input).
- `vllm/vllm-openai:v0.29.0` for the judge; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).
- Optional: `lfoppiano/grobid:0.8.2` (about 1.7 GB, CPU).
## 3. docker-compose.yml
Write this in `~/decosa/papercheck/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
grobid:
image: lfoppiano/grobid:0.8.2
ports: ["127.0.0.1:8070:8070"]
mem_limit: 8g
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_PAPERCHECK_MAILTO: "<contact address>"
DECOSA_PAPERCHECK_GROBID_URL: http://grobid:8070
DECOSA_PAPERCHECK_WORKERS: "6"
DECOSA_PAPERCHECK_MAX_CLAIMS: "300"
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy }, grobid: { condition: service_started } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/papercheck/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not a host folder:
the image runs as uid 10001, and a host folder Docker creates is root's, which stops the api with a PermissionError.
GROBID is optional: drop the service and `DECOSA_PAPERCHECK_GROBID_URL` to use the rule parser alone. To cache public
metadata across restarts, add `DECOSA_PAPERCHECK_CACHE_DIR: /data/cache`. Start: `docker compose up -d`.
On first start the api creates this box's Ed25519 key in the volume (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup`, keep it private and never print it. Every model call on the direct
route gets a receipt signed with it (status `attested`): an attestation by me, the operator, not a proof.
## 4. Smoke test
1. `curl -s localhost:8445/papercheck/info | jq '{route: .model.route, parser: .reference_parser, ua: .lookups.user_agent}'`
shows `direct`, `rules+grobid` (or `rules`) and my contact address.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"paper-claim-check"}' | jq -r .token)`.
3. The planted sample (a CC BY paper with three planted miscitations):
`curl -s -XPOST localhost:8445/papercheck/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"telecommuting-planted","stream":false}' > run.json`.
Expect reference 43 `retracted: true` (the 2020 hydroxychloroquine registry paper), the `[36]` super-spreader
sentence `unsupported` (a hospital Staphylococcus aureus paper), the `[42]` sentence `contradicted` with a quoted
passage, a `year_mismatch` on reference 1, and most other citations `supported`. The same citation can vary between
runs; report what you get.
4. Run the smoke module from the repository: `DECOSA_API_KEY=$T python scripts/smoke/paper-claim-check.py http://127.0.0.1:8445`
must print `"ok": true`.
5. `jq '{record: .report.record}' run.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
must show `ok: true`. Change one verdict in the record and verify again: it must fail.
6. `docker compose logs api | grep -ci hydroxychloroquine` must print 0: no manuscript text in the logs.
7. Time it: on our RTX PRO 6000 the 24-citation sample took 20-33 s with two calls in flight on a shared card. Tell me
what you measure.
## 5. Point your tools at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call the routes from your own
workflow: `POST /papercheck/extract` turns a PDF, DOCX or LaTeX file into text, `POST /papercheck/check` checks it, and
`GET /papercheck/runs/{id}/export?format=md` gives the editor's report. Contract: `API_CONTRACT.md`, section
"Citation and claim checker for papers".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds confidential
manuscripts. If I ask for it, follow the provider guide at `/provide` on the site, and only with my explicit yes.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Every citing sentence checked against the cited paper's open full text, with the passage quoted; retracted and corrected sources, wrong DOIs and years, and self-citation stacks flagged; one receipt per check and a signed report.
- Who it's for
- Teams in science and research.
- Where it runs
- Hosted for published papers and your own drafts; self-host for manuscripts under review
- Key numbers
- 68 of 68 / 68 of 68 Retraction flag on known retracted works, with DOI / without (held out, n = 68)
- 0 of 60 / 0 of 60 Retraction flag on random control articles, with DOI / without (held out, n = 60)
- 12 of 12 Planted wrong-paper citations flagged (test split, n = 12)
- 9.6 s Median end-to-end run, hosted (QA sweep 2026-09-25)
- Models
- Qwen3.8-27B
- Where
- Hosted for published papers and your own drafts; self-host for manuscripts under review
- Checks
- Receipt per citation checked; signed report
- Industry
- Science and research
- Output
- Structured data · Signed record or verdict
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Grounding · Signed record
Questions people ask
How does it decide whether a cited paper supports a sentence?
An open model reads the cited paper's open full text (PubMed Central open access through Europe PMC, or arXiv) and says supported, overstated, not supported or contradicted, quoting the passage. Citations without open full text are marked unavailable, never guessed.
Does it flag retracted references?
Yes. Each reference is looked up in Crossref and OpenAlex for retractions, withdrawals, expressions of concern and corrections; Crossref's notices include the Retraction Watch data. It also flags DOIs that do not exist or point to another paper, and wrong years or first authors.
How often is it wrong?
On six test papers it caught 12 of 12 planted wrong-paper citations and 12 of 12 overstatements, but flagged 12 of 49 correct citations (about one flag in four is noise). Labels were an AI agent's, not a domain expert's, and results can vary run to run.
Can a peer reviewer use it on a confidential manuscript?
Run the self-hosted version. Then only DOIs, identifiers and reference strings go to Crossref, OpenAlex, Europe PMC, arXiv and doi.org; the manuscript text stays on your machine. The hosted demo is for published papers or your own drafts.
Is a flag a finding of misconduct?
No. It is a screening aid. Citation manipulation, including self-citation stacking, is for the journal or institution to investigate, for example under COPE guidance.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Citation and claim checker for papers
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…