69 · Legal · Finance and insurance · live
M&A due-diligence red flags
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted red flags caught20 / 20test splitn = 20two held-out invented rooms (Osprey 11, Heron 9), wording set B, run once with prompts frozen
- Flags raised in a clean room of near misses0test splitheld-out room Wren
- Quotes found word for word at the cited byte span22 / 22test splitn = 22
- Planted flags caught in 40-document rooms (48 public contracts added)20 / 20test splitn = 20direct route; 0 flags raised in the added contracts
- CUAD anti-assignment clauses found9 / 12held outn = 1228 public CUAD contracts, precision 0.75; BM25 only: 11 / 12
- CUAD exclusivity and non-compete clauses found3 / 7held outn = 7precision 3/3; BM25 only: 4 / 7
- Planted red flags caught, dev11 / 11dev (tuned on)n = 11Kestrel; clean dev room Linnet: 0 flags
Dataset
Five invented data rooms of 16 documents (dev Kestrel and clean Linnet; held-out Osprey, Heron and clean Wren, worded from a second clause set written before any model run); the held-out rooms padded with 48 CUAD v1 contracts (CC BY 4.0); 28 further CUAD contracts scored against CUAD's expert labels.
Caveats
- Same author wrote the rooms, the gold and the prompts; the rooms are short and tidy compared with real data rooms.
- The invented rooms are at the ceiling: BM25 search and reading the whole room also found every planted flag, so they do not separate retrieval methods.
- On public CUAD contracts BM25 did as well or better than the reranked hybrid; small n (28 contracts).
- No real data room and no lawyer's review of the memos.
- Stress and CUAD runs used the direct route to the same model because the gateway was down.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 93 s
- Receipts
- 3
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.003
Self-host verification
Verified on 27 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running retrieval service, local signing; torn down after
The rehearsal bundle passed 13/13 in 101 s (Kestrel with its planted flags, the record verified and failed when changed, the clean Linnet room with no flags; every receipt attested). The retrieval service was not built from the compose file here: the running one was used.
Rehearsal bundle: ma-dd-redflags.zip (16 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- Measured on five small invented rooms written by the same author as the prompts, and 28 public contracts; not on a real data room or against a lawyer's issues list.
- Ten categories only; tax, employment, data protection, environmental and sanctions issues are not looked for.
- Text only: scanned documents must go through the document reader block first.
- 'No flag found' means nothing turned up in the passages the search returned for that category; a run where model calls failed is marked incomplete.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)AGPL-3.0-or-later
- One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each findingQwen3.8-27B (NVFP4)Apache-2.0
- Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · BM25 retrieval, no retrieval service (2)
- Planted flags caught, held-out rooms, BM25 only: 20 / 20, 0 in the clean roomdecosa-api docs/evals/ma-dd-redflags.md, test split (direct route), 27 Sep 2026
- CUAD public contracts: anti-assignment / exclusivity found, BM25 only: 11/12 (precision 0.917) / 4/7decosa-api docs/evals/ma-dd-redflags.md, CUAD check
Standard · retrieval service + the model (hosted demo) (5)
- Planted flags caught, held-out rooms (Osprey, Heron): 20 / 20decosa-api docs/evals/ma-dd-redflags.md, test split, 27 Sep 2026
- Flags raised in the clean held-out room (Wren): 0decosa-api docs/evals/ma-dd-redflags.md, test split
- Quotes found word for word at the cited byte span: 22 / 22 (held-out memo flags)decosa-api docs/evals/ma-dd-redflags.md, test split
- Planted flags caught with 48 public contracts added as distractors (40-document rooms): 20 / 20, 0 flags in the added contractsdecosa-api docs/evals/ma-dd-redflags.md, stress split (direct route)
- CUAD public contracts: anti-assignment / exclusivity / change of control found: 9/12 (precision 0.75) / 3/7 / 2/2decosa-api docs/evals/ma-dd-redflags.md, CUAD check, 28 contracts