Evidence retrieval
Finds the passages in your documents that answer a question, on open models: an embedder and BM25 pick candidates, a reranker orders them. Every hit is a chunk with byte offsets and a Merkle proof into a signed snapshot of the corpus, so a cited line can be traced to the exact document version it came from.
Measured 2026-09-27. Full eval. First uses: Security questionnaire, M&A red flags, tariff classification and medical chronology (scanned pages indexed with their boxes).
Watch real searches
Recorded runs on two sample corpora. Each hit is a chunk with its byte offsets in the source document. The receipt is signed; your browser re-checks the signature, the index hash and every chunk's Merkle proof, and a changed character fails.
Loading the recorded run…
How it works
- Documents are cut into chunks at paragraph and sentence boundaries (up to 1,200 characters, no overlap). Each chunk keeps its byte offsets in the source document and its SHA-256.
- The index hash commits to the corpus (a Merkle root over each chunk's document hash, id, offsets and hash), the chunker settings and the embedder's repo and revision. The same files give the same corpus root on any machine.
- Search: Qwen3-Embedding-0.6B (dense) and BM25 each propose 40 chunks, fused by reciprocal rank; Qwen3-Reranker-4B scores those 40 and orders them.
- Every embed and rerank call gets a decosa.model-call.v1 receipt: the model's weights root, hashes of the texts and of the vectors or scores, and the parameters.
- Every search returns a signed decosa.retrieval.search.v1: the index hash, the query hash, the reranker revision, and the chunks returned with their offsets, integer scores and ranks.
- Downstream, answers may only quote retrieved chunks: quotes are matched word for word back to a chunk and its bytes, and the grounding block judges each sentence against the retrieved chunks only.
A search receipt, shortened:
{
"v": "decosa.retrieval.search.v1",
"id": "rs-1f3ec45321c2af8ca415",
"index": {
"hash": "e10c3051bfd024bb…",
"corpus_root": "e398087f5b3e3b91…",
"n_chunks": 171,
"embedder": {
"repo": "Qwen/Qwen3-Embedding-0.6B",
"revision": "97b0c61"
}
},
"reranker": {
"repo": "Qwen/Qwen3-Reranker-4B",
"revision": "22e6836"
},
"query_sha256": "f25e731a9cbaea36…",
"params": {
"k": 5,
"rerank": true,
"mode": "hybrid",
"candidates": 40
},
"results": [
{
"rank": 1,
"chunk": "K1#20",
"doc": "K1",
"doc_sha256": "bc3db8c11b06d084…",
"byte_start": 16453,
"byte_end": 17515,
"sha256": "a33d10501ca38e5a…",
"scores_ppm": {
"dense": 449034,
"rerank": 294215
},
"ranks": {
"dense_rank": 14,
"bm25_rank": 1
}
},
"…"
],
"model_calls": [
{
"id": "mc-6b3baa596d807019067c",
"kind": "embed"
},
{
"id": "mc-d2e6078e2c42cb7889fe",
"kind": "rerank"
}
],
"signer": "700c585540114e5c…",
"sig": "8124f74db726b1d6…"
}API
| POST /retrieval/index | {documents: [{id, title, text}], synthetic: true} or {sample} -> the index (hash, corpus root, chunks with byte offsets and hashes) and the embedding call's receipt. Any valid session token or dk_ key. |
| POST /retrieval/search | {index, query, k?, rerank?, mode?: hybrid | dense | bm25} -> ranked chunks with text, byte offsets, scores and Merkle proofs, the signed search receipt and the model-call receipts. Documents can be sent here directly instead of an index. |
| POST /retrieval/index-scans | {documents: [{name, document}], synthetic: true} with documents read by POST /docreader/read -> an index whose chunks are page elements: every search result carries anchors (page, box, element) and the index has a layout hash over the boxes. Each document's reader receipt must verify and be signed by this server. |
| POST /retrieval/quotes | {index, query | chunks, answer | quotes} -> each quote found word for word in a retrieved chunk (with byte offsets) or not. No model call beyond the search. |
| POST /retrieval/verify | {receipt, chunks?, documents?} -> signature, index hash, and for each chunk its text hash, Merkle proof and (with the document) its offsets. No key. |
| GET /retrieval/info | Models and pinned revisions, weights roots, licences, limits, whether the model service is reachable, what it does not do. |
Hosted, the block takes synthetic or public documents only, up to 60 documents and 1,500 chunks per index, held in memory for an hour. Self-host it for real ones.
Which models, and why
Six open models compared on dev splits only: our security-questionnaire questions (26) and in-contract clause finding on CUAD (164 queries, 40 contracts). The test splits and two MLEB sets were scored once, after the choice. Memory is the measured peak on the card.
| Setup | Questionnaire dev: success@1 / nDCG@10 | CUAD dev: MRR@10 / recall@10 | GPU memory (peak) |
|---|---|---|---|
| BM25 (what we had) | 0.923 / 0.940 | 0.644 / 0.754 | none |
| Qwen3-Embedding-0.6B, dense | 0.846 / 0.909 | 0.763 / 0.880 | 3.0 GB |
| Nemotron-3-Embed-1B, dense | 0.885 / 0.885 | 0.781 / 0.864 | 3.8 GB |
| Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (chosen) | 1.000 / 1.000 | 0.753 / 0.879 | 3.0 + 9.6 GB |
| Nemotron-3-Embed-1B + Qwen3-Reranker-4B | 1.000 / 1.000 | 0.767 / 0.895 | 3.8 + 9.6 GB |
| Qwen3-Embedding-4B + Qwen3-Reranker-4B | 1.000 / 1.000 | 0.753 / 0.871 | 10.7 + 9.6 GB |
| Qwen3-Embedding-0.6B + Qwen3-Reranker-0.6B (lean) | 1.000 / 0.997 | 0.760 / 0.856 | 3.0 + 2.3 GB |
The pairs with the 4B reranker tie on the questionnaire; on CUAD the Nemotron embedder is 0.014 MRR ahead, within noise on 164 queries. We serve the all-Apache Qwen pair: one licence, one prompt family, and the use cases built on the block were measured on it. The larger Qwen3-Embedding-4B bought nothing on dev or held out and would put the service over its 18 GB share of the card. The lean pair is the choice for a box without the room.
Results
Held out: the questionnaire test split (64 questions), CUAD test (541 queries, 120 contracts) and two MLEB sets (Isaacus, CC BY 4.0) that had no dev split. Run once.
| Test | BM25 | Best embedder alone | Lean pair | Chosen pair |
|---|---|---|---|---|
| Questionnaire: success@1 | 0.875 | 0.969 (Qwen-4B) | 0.984 | 0.984 |
| Questionnaire: nDCG@10 | 0.925 | 0.957 (Qwen-4B) | 0.972 | 0.994 |
| MLEB contract clauses: success@1 | 0.533 | 0.867 (Nemotron-1B, Qwen-4B) | 0.822 | 0.933 |
| MLEB contract clauses: nDCG@10 | 0.613 | 0.905 (Qwen-4B) | 0.874 | 0.953 |
| MLEB Australian tax guidance: nDCG@10 | 0.556 | 0.733 (Qwen-4B) | 0.645 | 0.723 |
| CUAD clause in one contract: MRR@10 | 0.633 | 0.765 (Nemotron-1B) | 0.738 | 0.712 |
| CUAD clause in one contract: recall@10 | 0.775 | 0.880 (Nemotron-1B) | 0.865 | 0.881 |
Where the reranker does not help: finding a clause type inside one contract (CUAD), where the query is a long category description and the gold is an expert's highlight. The reranker finds more of the relevant chunks in the top 10 but ranks one first less often than the dense embedder alone. For that job, ask for mode dense without the reranker, or read the top 10.
| CUAD test (541) | success@1 | recall@10 | MRR@10 |
|---|---|---|---|
| Nemotron-1B dense, no reranker | 0.671 | 0.880 | 0.765 |
| Qwen-0.6B dense, no reranker | 0.614 | 0.860 | 0.728 |
| Chosen pair (hybrid + Qwen3-Reranker-4B) | 0.612 | 0.881 | 0.712 |
Measured on our server's shared GPU0. Most runs were taken while it was 89-97% busy with other jobs; the replay above was recorded with it quiet. The service runs one request at a time.
| What | Time | Notes |
|---|---|---|
| Embed one query (Qwen3-Embedding-0.6B) | 9 ms p50, 12 ms p95 | |
| Index chunks (up to 1,200 characters) | about 500 chunks a second | |
| Rerank 40 candidates (Qwen3-Reranker-4B) | 0.4-0.5 s with the GPU quiet; 1.8 s p50, 2.5 s p95 while it was 90% busy | the 0.6B reranker: 0.3 s p50 |
| A search through the API (index of 10-170 chunks) | 0.15-0.5 s quiet; 0.5-1.6 s under load | most of it is the rerank |
| Service memory (both models) | about 10 GB at rest; 11.8 GB peak with 4 clients, 12.4 GB seen under the use-case evals | ceiling 18 GB |
| On a Mac (M3 Ultra, PyTorch MPS) | embed 35 ms; rerank 40 in 5.1 s p50 | same scores within rounding: questionnaire nDCG 0.993, MLEB clauses 0.952, tax 0.718; 12.2 GB |
First upgrade: the security questionnaire
The security questionnaire answerer now takes its candidates from the block: 4 approved answers and 2 policy passages after reranking, instead of 6 and 3 by BM25. Prompts, the selector and the checks are unchanged. Same synthetic library and the same 100 held-out questions, run once.
| Measure | Before (BM25) | After (retrieval block) |
|---|---|---|
| Fill precision | 0.970 (64 of 66) | 0.984 (63 of 64) |
| Coverage of answerable questions | 0.984 (63 of 64) | 0.984 (63 of 64) |
| Abstention when nothing approved fits | 0.969 (31 of 32) | 1.000 (32 of 32) |
| Stale approved answers caught | 4 of 4 | 4 of 4 |
| Invented claims | 0 | 0 |
| Prompt tokens for the 100 questions | 218,203 | 151,150 (-31%) |
| Retrieval alone, no model call (top answer above a dev threshold): precision | 0.606 | 0.864 |
Retrieval alone now closes 71% of the gap between the old keyword baseline and the model's selection, with no model call, though it catches only 1 of 4 stale answers: the policy check still needs the model. On this small synthetic library the right source was always among the candidates either way, so the gains are a better top candidate and shorter prompts. Each question's search has a signed receipt, and the review record names the index hash.
Where it goes
Security questionnairebuilt
Candidates for the answer selector; a search receipt per question; the index hash in the signed review record.
M&A due-diligence red flagsbuilt
Each red-flag category searched across the data room; every flag quotes a retrieved chunk, checked word for word, cited to its byte span.
Tariff classification memobuilt
Similar CBP rulings and HTS text retrieved for a product description; the memo cites rulings by number with quotes checked against the retrieved text.
Medical chronologybuilt
Scans into retrieval: each page's document-reader elements become chunks that keep their page and box (a layout hash commits to the boxes); every chronology line's quote is located to its chunk, byte span and box, and a search per procedure finds other pages that give a different date.
Legal drafting, privilege log, deposition, patent claims, filing pre-flightnext
Replace each module's own keyword search for the lines a sentence cites.
Groundingnext
Evidence selection for long sources: the retrieved chunks, not BM25 windows, go to the judge.
Self-host
Two processes: the retrieval service holds the embedder and the reranker; decosa-api chunks, indexes, searches and signs.
uv venv .venv-retrieval --python 3.12
uv pip install -p .venv-retrieval -r services/retrieval/requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
hf download Qwen/Qwen3-Embedding-0.6B --revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3
hf download Qwen/Qwen3-Reranker-4B --revision 22e683669bc0f0bd69640a1354a6d0aebcfeede5
RETRIEVAL_EMBED=qwen3-emb-0.6b RETRIEVAL_RERANK=qwen3-rr-4b .venv-retrieval/bin/python services/retrieval/server.py # 127.0.0.1:8499
DECOSA_RETRIEVAL_URL=http://127.0.0.1:8499 DECOSA_RETRIEVAL_SYNTHETIC_ONLY=0 python -m decosa_apiThen rehearse on the bundled synthetic policies before any real document: python scripts/rehearse.py evidence-retrieval. About 12 GB of GPU memory for the pair we serve; the lean pair (RETRIEVAL_RERANK=qwen3-rr-0.6b) needs about 5 GB and runs on a CPU, slowly. On a Mac the service uses MPS by itself: measured on an M3 Ultra, same scores, reranking about 3 times slower than our GPU. The rehearsal indexes nine synthetic policies and checks the corpus root anyone gets from them.
Licences and receipts
| Qwen3-Embedding-0.6B | Apache-2.0, Qwen/Qwen3-Embedding-0.6B @ 97b0c61 |
| Qwen3-Reranker-4B | Apache-2.0, Qwen/Qwen3-Reranker-4B @ 22e6836 |
| Lean tier: Qwen3-Reranker-0.6B | Apache-2.0, @ e61197e |
| Measured, not served | Nemotron-3-Embed-1B (OpenMDW-1.1), llama-nemotron-rerank-1b-v2 (OpenMDW-1.1 with the Llama 3.2 notice), Qwen3-Embedding-4B (Apache-2.0) |
| Eval data | CUAD v1 and two MLEB sets, all CC BY 4.0; our synthetic questionnaire library |
| Not used | Jina embeddings and rerankers v3.5+ and ContextualAI rerankers (non-commercial licences); MLEB sets under CC BY-NC |
The embedder and reranker run outside our gateway, so their receipts are signed by decosa-api's own key (an attestation by its operator) and, where the gateway is configured, reported to it and countersigned as a witness of the hashes, not the bytes. The search receipt is signed the same way. What they prove: which corpus snapshot and model revisions were used and which chunks came back with which scores. What they do not: that the ranking is right.
What it does not do
- English only: no other language was measured.
- Finding one clause type inside one long contract: the reranker does not help (see above); use dense mode or read the top 10.
- The hosted index lives in memory for an hour, for synthetic or public documents only; it is not a database. Self-host for a persistent corpus of real documents.
- Scanned PDFs are not read here: pass the text, or the document reader's elements.
- A high reranker score means relevant, not correct: whether a passage supports an answer is the grounding block's job.