Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Evidence retrieval block: eval (27 Sep 2026)

Questions:

  • Does an open embedder plus reranker find the passage that answers a question better than BM25?
  • Which pair should the block serve on GPU0 within its 18 GB ceiling?
  • Does it make a live use case better?

Code: decosa_api/retrieval/ (block), services/retrieval/ (models), scripts/retrieval_eval.py (this eval). Results: docs/evals/retrieval/{dev,test}.json (every config, every metric) and docs/evals/retrieval/qn-top1.json. Model timings and memory are in <internal path> on our server.

What was compared

First stage.

  • BM25: the grounding block's, which the questionnaire used before.
  • Dense cosine over each embedder's vectors.
  • Hybrid: dense and BM25 fused by reciprocal rank, k = 60.

Second stage. A reranker scores the top 40 first-stage chunks and orders them.

Model Licence (checked 27 Sep 2026 on the HF card and LICENSE) Revision
Qwen/Qwen3-Embedding-0.6B Apache-2.0 97b0c61
nvidia/Nemotron-3-Embed-1B-BF16 OpenMDW-1.1 ("ready for commercial use") c0c9fea
Qwen/Qwen3-Embedding-4B Apache-2.0 5cf2132
Qwen/Qwen3-Reranker-0.6B Apache-2.0 e61197e
nvidia/llama-nemotron-rerank-1b-v2 OpenMDW-1.1, plus the Llama 3.2 Community Licence notice ("Built with Llama"); the repo ships its own model code (trust_remote_code, pinned) 8287656
Qwen/Qwen3-Reranker-4B Apache-2.0 22e6836
  • All models ran in bf16 on GPU0 (RTX PRO 6000) with transformers 5.17 and torch 2.14.
  • Each model was loaded alone, and memory was freed between models.
  • Each model gets the card's own prompt format:
    • the Qwen embedders get a task instruction on queries only;
    • the Qwen rerankers use the yes/no prompt, scored as P(yes);
    • Nemotron uses its sentence-transformers prompts and its question:/passage: template.

Chunking.

  • Paragraphs are merged up to 1,200 characters and cut at sentence ends.
  • Questionnaire policies use 700 characters, as the questionnaire has always done.
  • Chunks do not overlap.

Data (all licence-clean)

Set Licence What Queries
qn: our security-questionnaire split ours (synthetic) 50 approved answers + 9 policies of the fictional Tallyloom; the gold is the answer id or the policy id (written 25 Sep, before any run) dev 26, test 64 (the fill questions)
cuad: CUAD v1 clause finding CC BY 4.0 (The Atticus Project); the contracts are public SEC EDGAR exhibits Find the clause in one contract, as a data-room review does. 18 deal-relevant categories were fixed before any run (change of control, anti-assignment, exclusivity, non-compete, termination for convenience, IP assignment, joint IP, MFN, uncapped liability, cap on liability, minimum commitment, liquidated damages, ROFR/ROFO/ROFN, revenue sharing, audit rights, insurance, non-transferable licence, covenant not to sue). A chunk is relevant when it overlaps an expert answer span. Contracts are split by sha256(title): dev 40, test 120, each at most 120k characters. dev 164, test 541
mleb-clauses: MLEB Contractual Clause Retrieval (Isaacus) CC BY 4.0 45 clause definitions → 90 example clauses 45 (no dev split)
mleb-tax: MLEB Australian Tax Guidance Retrieval (Isaacus) CC BY 4.0 112 real tax questions → 105 ATO guidance pages 112 (no dev split)
  • Other MLEB sets were left out on licence: Consumer Contracts QA and Legal RAG Bench (CC BY-NC), GDPR Holdings (CC BY-NC-SA).
  • Unit:
    • cuad is scored per chunk;
    • qn and mleb are scored per document, ranked at its best chunk.
  • Metrics: success@k (any relevant unit in the top k), recall@10, MRR@10, nDCG@10, all with binary relevance.

How the pair was chosen (dev only)

  • The ranking was frozen on the qn and cuad dev splits.
  • The test splits and both MLEB sets were scored once, after the choice.
  • Nothing was tuned on them: no prompt, chunk size, fusion constant or candidate count was changed after seeing them.
Dev BM25 Qwen-0.6B dense Nemotron-1B dense Qwen-0.6B hybrid + Qwen-Reranker-4B Nemotron-1B hybrid + Qwen-Reranker-4B Qwen-4B hybrid + Qwen-Reranker-4B Qwen-0.6B + Qwen-Reranker-0.6B
qn success@1 / nDCG@10 0.923 / 0.940 0.846 / 0.909 0.885 / 0.885 1.000 / 1.000 1.000 / 1.000 1.000 / 1.000 1.000 / 0.997
cuad MRR@10 / recall@10 0.644 / 0.754 0.763 / 0.880 0.781 / 0.864 0.753 / 0.879 0.767 / 0.895 0.753 / 0.871 0.760 / 0.856
GPU memory, peak (measured) 0 3.0 GB 3.8 GB 3.0 + 9.6 GB 3.8 + 9.6 GB 10.7 + 9.6 GB 3.0 + 2.3 GB

Choice: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B, hybrid first stage, 40 candidates. Why:

  • Every reranker pair reaches 1.000 success@1 on qn dev. The 4B reranker is the only one at nDCG 1.000.
  • On cuad dev, the Nemotron-1B embedder is ahead of Qwen-0.6B by 0.014 MRR over 164 queries. That is within noise; the two pairs are tied on qn.
  • We took the all-Apache Qwen pair. It is one licence and one prompt family, and the two use cases on top of the block (M&A red flags, tariff classification) were measured on it. Nemotron-1B + Qwen-Reranker-4B is the documented alternative.
  • The "quality tier" page 48 proposed (Qwen-4B + Qwen-Reranker-4B) peaked at about 20 GB when measured, over the 18 GB ceiling. It was no better on dev (below), and no better held out either.
  • The lean tier is Qwen3-Embedding-0.6B + Qwen3-Reranker-0.6B (about 5 GB; CPU-capable), for boxes without the room.

Held-out results (test splits and MLEB, run once)

Test BM25 Qwen-0.6B dense Nemotron-1B dense Qwen-4B dense Lean: Qwen-0.6B + Rerank-0.6B Chosen: Qwen-0.6B + Rerank-4B Qwen-4B + Rerank-4B
qn (64): success@1 0.875 0.953 0.938 0.969 0.984 0.984 0.984
qn: nDCG@10 0.925 0.944 0.941 0.957 0.972 0.994 0.994
cuad (541): MRR@10 0.633 0.728 0.765 0.742 0.738 0.712 0.717
cuad: recall@10 0.775 0.860 0.880 0.862 0.865 0.881 0.886
mleb-clauses (45): success@1 0.533 0.800 0.867 0.867 0.822 0.933 0.933
mleb-clauses: nDCG@10 0.613 0.852 0.886 0.905 0.874 0.953 0.953
mleb-tax (112): MRR@10 0.493 0.583 0.623 0.669 0.568 0.662 0.664
mleb-tax: nDCG@10 0.556 0.660 0.689 0.733 0.645 0.723 0.732

All 19 configurations with every metric are in docs/evals/retrieval/test.json.

Reading it:

  • Every learned setup beats BM25 everywhere. On MLEB clauses, success@1 goes from 0.533 to 0.933. On tax, nDCG goes from 0.556 to 0.723. On in-contract clause finding, MRR goes from 0.633 to 0.71–0.77.
  • The 4B reranker is where the held-out gain on MLEB comes from. With the 0.6B reranker, clauses nDCG is 0.874 and tax 0.645; with the 4B reranker it is 0.953 and 0.723.
  • The reranker does not help in-contract clause finding (cuad).
    • It raises recall@10 (0.860 → 0.881) but lowers top-1 (MRR 0.728 → 0.712).
    • Nemotron-1B dense with no reranker is best there (MRR 0.765).
    • CUAD queries are long category descriptions, and relevant chunks are partial expert highlights.
    • Use cases that scan one document per clause type should ask for rerank: false, mode: "dense", or take the top 10 rather than the top 1.
    • The M&A red-flag use case measured the same thing on its CUAD check (see docs/evals/ma-dd-redflags.md).
  • The bigger embedder is not worth the memory.
    • Qwen-4B + Reranker-4B equals the chosen pair within 0.01 on every held-out set.
    • Qwen-4B dense alone is best on tax (0.733 nDCG), which the reranked pairs match, not beat.
  • The benchmark claim to compare with. Page 32 quotes Qwen3-Embedding-8B at MLEB 0.759 overall. Our two MLEB sets are a subset, scored on our own chunking, so our numbers are not comparable to the leaderboard.

Latency and memory (measured on our server, GPU0 shared, 27 Sep 2026)

GPU0 was 89–97% busy with other jobs (Voxtral, the language-pack build, vLLM, the prover) through every run, so these timings are under load.

Qwen-0.6B embed Nemotron-1B embed Qwen-4B embed Qwen-Reranker-0.6B Nemotron rerank-1B Qwen-Reranker-4B
One query embedded, p50 / p95 9 / 12 ms 6 / 9 ms 14 / 18 ms
Indexing throughput (chunks ≤ 1,200 chars) 496 / s 386 / s 105 / s
Rerank 40 candidates, p50 / p95 316 / 546 ms 135 / 216 ms 1,758 / 2,494 ms
Peak GPU memory (torch, batch 16 embed / 8 rerank) 3.0 GB 3.8 GB 10.7 GB 2.3 GB 2.8 GB 9.6 GB

The service as deployed (decosa-retrieval: Qwen-0.6B + Qwen-Reranker-4B).

  • Memory: 10.3 GB allocated at rest (the weights). The process peaked at 11.8 GB (nvidia-smi) under 4 concurrent clients, each sending a 40-candidate rerank and a 256-chunk embed three times. Later, under the two use cases' eval runs, the torch peak reached 12.4 GB. Both are inside the 18 GB ceiling.
  • Rerank latency in that test: 1.0–1.3 s per 40 candidates, up to 6 s at the back of the queue (requests run one at a time).
  • End to end through decosa-api: 0.15–0.5 s for a search of a 10–171 chunk corpus with GPU0 quiet (the site replay, 27 Sep 12:52 UTC; the 40-candidate rerank took 0.14–0.48 s of it), 0.5–1.6 s under load.
  • A bug fixed on the way: the Qwen rerankers first computed logits for every position, about 4 GB extra at batch 8. The service now keeps the last position only, and it frees cached memory after each call.

First upgrade: the security questionnaire (before → after)

  • What changed: the answerer's candidates now come from the block (4 approved answers + 2 policy passages after reranking) instead of BM25 (6 + 3).
  • What did not change: the prompts, the selector and the checks.
  • Receipts: each question gets a signed search receipt, and the review record names the index hash.
  • The test split: run once, after dev.
Held-out test (100 questions) Before (BM25 candidates, 25 Sep) After (retrieval block, 27 Sep)
Fill precision 0.970 (64 of 66) 0.984 (63 of 64)
Coverage of answerable questions 0.984 (63 of 64) 0.984 (63 of 64)
Abstention when nothing approved fits 0.969 (31 of 32) 1.000 (32 of 32)
Stale approved answers caught 4 of 4 4 of 4
Invented claims (lexical check) 0 (0 misses) 0 (0 misses)
Model calls / prompt tokens / generated tokens 177 / 218,203 / 11,695 172 / 151,150 (−31%) / 9,003
Dev (40): precision / coverage / abstention / stale 0.963 / 1.000 / 0.917 / 2 of 2 0.963 / 1.000 / 0.917 / 2 of 2 (prompt tokens 79,741 → 60,372)

What changed on test:

  • The HSM question (t084), filled wrongly before, now goes to a person.
  • The one remaining error is t056, the arguable laptop-lock label already discussed in the questionnaire eval.
  • Wall time was 693 s against 548 s, but the gateway was far busier on 27 Sep. It is not a like-for-like speed comparison.

Page 32's question (how much of the 0.606 → 0.970 gap does the reranker close alone, with no model call).

  • The same top-1 baseline was run: the top approved answer is filled when its score clears a threshold chosen on dev.
  • With the block's reranker score instead of BM25, test fill precision is 0.864, coverage 0.797 and abstention 0.844 (threshold 0.47 on P(yes)).
  • BM25 top-1 scored 0.606 / 0.672 / 0.625.
  • So the reranker alone closes 71% of the precision gap.
  • It catches 1 of 4 stale answers. The policy check, which needs the model, still matters.

Candidate recall is not the bottleneck on this synthetic library. The gold was among the selector's candidates for every fill question in both regimes (BM25 6 + 3, and the block at 4 + 2, 3 + 1 or 2 + 1). The gain here is shorter prompts and a better top candidate. On real libraries with hundreds of overlapping answers, recall should matter more; that is unmeasured.

Receipts

  • Model-call receipts. Every embed and rerank call gets a decosa.model-call.v1 receipt (decosa_api/model_call.py):
    • the use case, the model and its hf-files-v1 weights root (computed by the service from the bytes of the snapshot it loaded, with the same code as scripts/model_root.py, and equal to the root recomputed from the Hub: 0.6B embedder 6f22380b…, 4B reranker 73b8e9c2…). Correction, 28 Sep 2026: until then the service and the script took Hugging Face cache blob names as SHA-256s; this cache was downloaded through Xet, whose blob names are Xet hashes, so receipts from 27 to 28 Sep carry wrong roots (embedder 2589cfee…, reranker c0e505a3…). Their signatures still verify; the root field in them is wrong. See decosa_api/weights_root.py;
    • hashes of the input texts and the output vectors or scores;
    • the parameters and the item count.
    • It is signed by decosa-api's key and, when the gateway is configured, witnessed by it ("reported": the gateway sees hashes, not bytes).
  • Search receipts. Every search returns a signed decosa.retrieval.search.v1:
    • the index hash: sha256 over the corpus Merkle root, the chunk count, the chunker, and the embedder repo and revision;
    • the query hash, the reranker revision, and the result chunks with byte offsets and integer scores (ppm);
    • the model-call receipt ids.
  • Checking a cited chunk. Its text, document hash and Merkle proof recompute the corpus root without the rest of the corpus. POST /retrieval/verify and the site's browser verifier both check this.
  • The rehearsal. It checks that the nine Tallyloom policies give corpus root d5c40972… on any machine.

Caveats

  • Small or synthetic sets. The questionnaire split is synthetic and single-author, with 26 dev and 64 test queries; every learned setup is near the ceiling on it.
  • CUAD relevance is partial. It is chunk overlap with expert highlights, which are often fragments. A chunk that holds the clause but not the highlighted words counts as a miss.
  • MLEB is not the full benchmark. It is two of its ten sets, on our chunking, so it is not comparable to the leaderboard.
  • English only. No other language was measured.
  • Latency is under load. GPU0 was shared and busy throughout; the timings are for that state, not a quiet card.
  • The dev choice was close. Nemotron-1B + Reranker-4B was within 0.014 MRR on cuad dev. The deciding factors (one licence, and continuity with the two use cases' evals) are judgement, not measurement.
  • What a reranker score is. It is a model's relevance judgement, not a check that the passage answers the question. The grounding block (17) and the quote check do that.

Reproduce (on our server)

export HF_HUB_CACHE=<internal path> HF_HUB_OFFLINE=1
PY=<internal path>
$PY scripts/retrieval_eval.py models                 # embeds and reranks every task, one model at a time (cached)
$PY scripts/retrieval_eval.py score --split dev      # -> docs/evals/retrieval/dev.json
$PY scripts/retrieval_eval.py score --split test     # -> docs/evals/retrieval/test.json (qn/cuad test + both MLEB sets)
$PY scripts/retrieval_eval.py qn-top1                # questionnaire top-1 baseline + candidate recall
# questionnaire before/after (gateway; retrieval service on :8499)
python scripts/questionnaire_eval.py run   --split test --work W --retrieval http://127.0.0.1:8499
python scripts/questionnaire_eval.py score --split test --work W

Data sources:

  • isaacus/contractual-clause-retrieval @ 3703358;
  • isaacus/australian-tax-guidance-retrieval @ 982aa11;
  • theatticusproject/cuad @ a3c393f (CUAD_v1/CUAD_v1.json).

All three are under <internal path>. The questionnaire runs are in docs/evals/security-questionnaire/retrieval/.

On a Mac (M3 Ultra, 512 GB; PyTorch MPS, bf16; 27 Sep 2026)

  • The code: the same service code (services/retrieval/models.py, which now picks MPS on Apple silicon) and the same served pair. The script is scripts/retrieval_eval_mac.py; the results are in docs/evals/retrieval/mac.json.
  • The sets: the held-out qn test and both MLEB sets, with the chosen setup only.
Mac (MPS) Our server (CUDA)
qn test nDCG@10 / success@1 0.993 / 0.984 0.994 / 0.984
MLEB clauses nDCG@10 / success@1 0.952 / 0.933 0.953 / 0.933
MLEB tax nDCG@10 / MRR@10 0.718 / 0.658 0.723 / 0.662
Embed one query, p50 35 ms 9 ms (under load)
Rerank 40 candidates, p50 / p95 5.1 / 8.5 s 1.8 / 2.5 s (under load)
Peak memory 12.2 GB (MPS driver) 11.8 GB (GPU, service)

Reading it:

  • Quality: the same within bf16 rounding.
  • Speed: the 4B reranker is about 3 times slower on MPS. The Mac tier should use the 0.6B reranker (RETRIEVAL_RERANK=qwen3-rr-0.6b) where interactive speed matters (not measured on the Mac).