M&A due-diligence red flags: eval (27 Sep 2026)
Questions: does the review find the red flags planted in a data room, with quotes that are really there and in the right place; does it stay quiet on a clean room full of near misses; and does the evidence retrieval block earn its place against plain keyword search or reading the whole room?
Data
- Invented data rooms (
decosa_api/verticals/madd/data/rooms.json, built byscripts/madd_rooms.py, Apache-2.0). Five rooms of 16 documents each: a customer MSA, an inbound software licence, a reseller agreement, a supply agreement, a public-sector services agreement, a credit agreement, a contractor agreement, the IP assignment register, the CEO's employment agreement, an office lease, an NDA, two sets of board minutes, the disclosure schedule, revenue by customer and a financial summary. About 22,000-24,000 characters per room. - Each room is assembled from a clause library in which every slot has a flagged and a clean wording. Clean wordings are near misses on purpose: a change of control that needs only notice, assignment allowed to a successor, a non-exclusive appointment, a capped indemnity, a covenant met with headroom, a contractor who "hereby assigns", a lawsuit the disclosure schedule lists, a largest customer at 9%.
- Two independent wording sets: set A for the dev rooms, set B (written separately, never seen while the prompts,
queries and definitions were tuned) for the held-out rooms.
- dev: Kestrel (11 planted flags), Linnet (clean).
- test (held out): Osprey (11), Heron (9), Wren (clean). Osprey and Heron share set B's wording, with different companies, numbers and mixes of flags.
- Gold (the planted flag, its type and its byte spans) is written by the generator from where it placed each flagged wording, before any model run (commit c95f100). One flag can live in two places (the contractor clause and the IP register line): either span counts.
- Stress rooms: Osprey and Heron each padded with 24 public contracts from CUAD v1 (CC BY 4.0), to 40 documents, about 270,000-290,000 characters and 514-563 chunks: the hosted maximum. CUAD's contracts are between other companies, so they are distractors, not the target's contracts.
- CUAD clause check: 28 public CUAD contracts (5,000-15,000 characters, not used in the stress rooms), one per room, "the target" read as either party, five categories that map to CUAD's expert labels. Scored against CUAD's labels, run once with prompts frozen.
Metrics
- Caught: a planted flag counts as caught when any flag in the memo quotes bytes that overlap one of its gold spans (any category); category right when that flag's category (or a merged one) is the gold type.
- False flags: any flag in a clean room; unplanted flags: flags in a planted room that overlap no gold span (listed for reading, none occurred).
- Quotes: every memo flag's quote must be found word for word at its byte span (code; a quote found nowhere is dropped and counted).
- CUAD: per contract and category, flagged or not against CUAD's label (precision, recall), and whether the flag's quote overlaps a CUAD answer span of that type.
Results
Standard tier (retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B, hybrid search; Qwen3.8-27B NVFP4, temperature 0, thinking off; grounding on). Baselines use the same prompts: BM25 (keyword search only, no retrieval service) and all (no retrieval: every chunk of the room in each category call).
| Split | Retrieval | Route | Planted caught | Category right | Flags in clean room | Unplanted flags | Quotes dropped | Model calls | Prompt / generated tokens |
|---|---|---|---|---|---|---|---|---|---|
| Test (held out) | hybrid + rerank | gateway | 20 / 20 | 20 / 20 | 0 (Wren) | 0 | 0 | 52 | 64,461 / 5,618 |
| Test | BM25 | direct | 20 / 20 | 20 / 20 | 0 | 0 | 0 | 52 | 69,802 / 6,023 |
| Test | all (no retrieval, no grounding) | direct | 20 / 20 | 20 / 20 | 0 | 0 | 0 | 30 | 193,509 / 4,855 |
| Stress (test + 48 CUAD contracts) | hybrid + rerank | direct | 20 / 20 | 20 / 20 | n/a | 0 (0 in CUAD contracts) | 0 | 42 | 44,803 / 5,188 |
| Stress | BM25 | direct | 20 / 20 | 20 / 20 | n/a | 0 | 0 | 42 | 53,138 / 5,483 |
| Dev | hybrid + rerank | gateway | 11 / 11 | 11 / 11 | 0 (Linnet) | 0 | 0 | 32 | 41,037 / 3,119 |
| Dev | BM25 | gateway | 11 / 11 | 11 / 11 | 0 | 0 | 0 | 31 | 43,259 / 2,931 |
Memo flags: 22 on test for 20 planted flags; the two extra flags are second quotes of a caught flag (the IP gap quoted from both the contractor agreement and the register), both overlapping gold. Every one of the 22 quotes was found word for word at its byte span. The grounding judge found every finding supported except one per planted test room, which the memo marks "check" for a person (Osprey 1, Heron 1).
CUAD clause check (28 public contracts, held out, no grounding, direct route):
| Category (CUAD label) | Hybrid: precision / recall | BM25: precision / recall | Contracts with the label |
|---|---|---|---|
| Change of control | 2/2 / 2/2 | 2/2 / 2/2 | 2 |
| Anti-assignment | 9/12 (0.75) / 9/12 (0.75) | 11/12 (0.917) / 11/12 (0.917) | 12 |
| Exclusivity, non-compete | 3/3 / 3/7 (0.429) | 4/5 (0.8) / 4/7 (0.571) | 7 |
| MFN | no flags, no labels | no flags, no labels | 0 |
| Uncapped liability | 0/4 / n/a | 0/2 / n/a | 0 |
The four "uncapped liability" flags quote indemnities in contracts that have no liability cap at all; CUAD labels none of them "Uncapped Liability". A reviewer might want them, but they are scored as false.
Latency (per 16-document room, 10 categories, grounding on): 21-35 s on the shared gateway on a normal day (Osprey 34.5 s, Heron 32.1 s, Wren 20.8 s); 219 s for Kestrel while the gateway was heavily loaded; 2-11 s on the direct route for the BM25 and no-retrieval baselines; about 50 s per room in the self-host check with the retrieval service (run while other evals shared the GPUs). A 40-document stress room: 47-58 s on the direct route (indexing about 550 chunks and 25 reranked searches). The gateway was down for about 15 minutes during this session; the eval runner now reruns a room whose categories failed, and a run with failed categories is marked incomplete in the memo and the record rather than reported as clean.
Self-host check: a fresh clone at 1656f11, the API image built from docker/api/Dockerfile, compose api with a
named volume, direct route to the local Qwen3.8-27B, the running retrieval service, local signing,
DECOSA_MADD_SYNTHETIC_ONLY=0: the rehearsal bundle passed 13/13 (101 s), then torn down.
Reading the numbers
- The synthetic rooms are at the ceiling. Every configuration, including plain BM25 and no retrieval at all, finds all 31 planted flags. On rooms this small and this tidy the model does the work, and retrieval cannot be told apart. These numbers show the mechanism works (search, one call per category, quote gate, grounding, record); they are not an estimate of accuracy on real data rooms.
- What retrieval buys here is size and proof, not recall. Reading the whole 16-document room costs about three times the prompt tokens (193,509 against 64,461 on test), and a real data room of hundreds of documents does not fit in one call at all; retrieval keeps each call to the eight best passages (about 1,300 prompt tokens per category on the 550-chunk stress rooms). The index hash and the signed search receipts say which corpus version and which passages the model saw.
- On public contracts, BM25 did as well or better than the reranked hybrid (anti-assignment 11/12 against 9/12; exclusivity 4/7 against 3/7). One contract per room means few chunks, so keyword queries that name the clause work; the dense side did not help. Small n (28 contracts); we did not change the default on this evidence, and the lite tier (BM25, no retrieval service) is a fair choice for small rooms.
- Nothing invented reached a memo: 0 quotes dropped on test, and the quote gate is covered by unit tests (an invented quote is dropped; a quote cited to the wrong excerpt is re-cited).
- Recall on public contracts is the real weak spot: exclusivity and non-compete clauses were missed in 3-4 of 7 contracts. The prompt's "affiliates" framing (restrictions that bind the buyer after closing) is narrower than CUAD's labels, which include any exclusivity.
Honest caveats
- One author wrote the rooms, the gold, the queries and the prompts. The held-out rooms use a separately written wording set, but the same author's idea of what a clause looks like.
- The rooms are short, clean text: no scans, no exhibits, no amendments that override earlier terms, no cross-references between documents beyond the minutes and the schedule.
- 31 planted flags and 28 public contracts is small n. No lawyer reviewed the memos.
- The grounding judge is a model judgement (see
docs/evals/grounding.md); here it reads one short finding against one passage, its easiest case. - The stress and CUAD runs used the direct route to the same Qwen3.8-27B server (the gateway was down); the held-out headline run used the gateway.
Expected properties of the sample run (Kestrel, rehearsal/ma-dd-redflags/)
- A change-of-control flag in the customer MSA (
MSA-BRI). - An undisclosed-litigation flag quoting the March board minutes (
MIN-2026Q1), which the disclosure schedule omits. - A debt-covenant flag whose quote contains
3.42x. - A customer-concentration flag from the revenue table (
REV-FY2025). - No flag in the plain NDA; the signed record verifies and fails when changed; the clean room Linnet gives at most one flag.
Verdict
Would a buyer pay? For a self-hosted first-pass triage on NDA-bound rooms, plausibly yes: the quote-and-byte-span discipline and the signed chain from index to memo are what a deal team needs to trust and defend a machine first pass, and nothing leaves the firm's box. What is missing before a paid pilot: a real, messy data room (hundreds of documents, scans through the document reader, amendments), measured recall against a lawyer's issues list, more categories (tax, employment, data protection, environmental), wider exclusivity recall, and a reviewer workflow (accept or reject each flag, then sign).
Reproduce
python scripts/madd_rooms.py # rebuilds rooms.json and the gold
python scripts/madd_eval.py run --split dev --mode hybrid # model calls; --mode bm25 | all
python scripts/madd_eval.py run --split test --mode hybrid # held out: run once
python scripts/madd_eval.py run --split stress --mode hybrid # test rooms + 48 CUAD contracts
python scripts/madd_eval.py cuad --mode hybrid # CUAD clause check
Runs and scores: docs/evals/ma-dd-redflags/*.json. CUAD v1 is read from
<internal path> (from theatticusproject/cuad @ a3c393f).