Skip to content
decosa

69 · Legal · Finance and insurance · live

M&A due-diligence red flags

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Planted red flags caught20 / 20test splitn = 20two held-out invented rooms (Osprey 11, Heron 9), wording set B, run once with prompts frozen
  • Flags raised in a clean room of near misses0test splitheld-out room Wren
  • Quotes found word for word at the cited byte span22 / 22test splitn = 22
  • Planted flags caught in 40-document rooms (48 public contracts added)20 / 20test splitn = 20direct route; 0 flags raised in the added contracts
  • CUAD anti-assignment clauses found9 / 12held outn = 1228 public CUAD contracts, precision 0.75; BM25 only: 11 / 12
  • CUAD exclusivity and non-compete clauses found3 / 7held outn = 7precision 3/3; BM25 only: 4 / 7
  • Planted red flags caught, dev11 / 11dev (tuned on)n = 11Kestrel; clean dev room Linnet: 0 flags

Dataset

Five invented data rooms of 16 documents (dev Kestrel and clean Linnet; held-out Osprey, Heron and clean Wren, worded from a second clause set written before any model run); the held-out rooms padded with 48 CUAD v1 contracts (CC BY 4.0); 28 further CUAD contracts scored against CUAD's expert labels.

Caveats

  • Same author wrote the rooms, the gold and the prompts; the rooms are short and tidy compared with real data rooms.
  • The invented rooms are at the ceiling: BM25 search and reading the whole room also found every planted flag, so they do not separate retrieval methods.
  • On public CUAD contracts BM25 did as well or better than the reranked hybrid; small n (28 contracts).
  • No real data room and no lawyer's review of the memos.
  • Stress and CUAD runs used the direct route to the same model because the gateway was down.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
27 Sep 2026
Latency, this run
n/a
p50 over passed runs
93 s
Receipts
3
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 27 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running retrieval service, local signing; torn down after

The rehearsal bundle passed 13/13 in 101 s (Kestrel with its planted flags, the record verified and failed when changed, the clean Linnet room with no flags; every receipt attested). The retrieval service was not built from the compose file here: the running one was used.

Rehearsal bundle: ma-dd-redflags.zip (16 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
  • Measured on five small invented rooms written by the same author as the prompts, and 28 public contracts; not on a real data room or against a lawyer's issues list.
  • Ten categories only; tax, employment, data protection, environmental and sanctions issues are not looked for.
  • Text only: scanned documents must go through the document reader block first.
  • 'No flag found' means nothing turned up in the passages the search returned for that category; a run where model calls failed is marked incomplete.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)AGPL-3.0-or-later
  • One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each findingQwen3.8-27B (NVFP4)Apache-2.0
  • Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · BM25 retrieval, no retrieval service (2)
  • Planted flags caught, held-out rooms, BM25 only: 20 / 20, 0 in the clean roomdecosa-api docs/evals/ma-dd-redflags.md, test split (direct route), 27 Sep 2026
  • CUAD public contracts: anti-assignment / exclusivity found, BM25 only: 11/12 (precision 0.917) / 4/7decosa-api docs/evals/ma-dd-redflags.md, CUAD check
Standard · retrieval service + the model (hosted demo) (5)
  • Planted flags caught, held-out rooms (Osprey, Heron): 20 / 20decosa-api docs/evals/ma-dd-redflags.md, test split, 27 Sep 2026
  • Flags raised in the clean held-out room (Wren): 0decosa-api docs/evals/ma-dd-redflags.md, test split
  • Quotes found word for word at the cited byte span: 22 / 22 (held-out memo flags)decosa-api docs/evals/ma-dd-redflags.md, test split
  • Planted flags caught with 48 public contracts added as distractors (40-document rooms): 20 / 20, 0 flags in the added contractsdecosa-api docs/evals/ma-dd-redflags.md, stress split (direct route)
  • CUAD public contracts: anti-assignment / exclusivity / change of control found: 9/12 (precision 0.75) / 3/7 / 2/2decosa-api docs/evals/ma-dd-redflags.md, CUAD check, 28 contracts

How we measure · All tools