Write a due-diligence red-flag memo
A red-flag memo for counsel where every flag quotes the exact words from the data room, with the document, severity and what to ask the seller.
Built on: Evidence retrieval, Grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B plus about 10 GB for the embedder and reranker; the quote checks, memo and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the m&a due-diligence red flags API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- ma-dd-redflags
Use the hosted API
# Decosa M&A due-diligence red flags: use the hosted API
You are wiring Decosa's due-diligence red-flag review into this project (a deal-room tool, a diligence tracker or a
script). It takes a data room as text and returns a red-flag memo for counsel: each flag with its category, severity,
the document, the exact words quoted, the chunk and byte span, why it matters and what to ask the seller, plus a signed
record. Every model call has a signed receipt and every search has a signed receipt naming the index hash. Use only what
is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes invented data rooms only** (`"synthetic": true` is required). Real data rooms are under NDA:
use the self-hosted version for them. Never send a real client document here.
- This is a triage aid for counsel, not legal advice. Never present its memo as a clean bill of health: "no flag found"
means nothing turned up in the passages searched.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "ma-dd-redflags"}` returns `{"token", "expires_at", "budget"}`.
The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response). Over a limit you get HTTP 429
with `Retry-After`. A demo token runs one review at a time (409 otherwise).
3. A full review needs about 7,500 generated tokens left in the budget (402 otherwise); fewer categories need less.
## Review a room
- `POST /madd/review` (token). Body: `{"room": {"target": "Example Widgets, Inc.", "buyer"?: "...", "title"?: "...",
"documents": [{"id"?, "title"?, "text"}]}, "synthetic": true, "categories"?: [...], "ground"?: true, "stream"?: false}`
or `{"sample": "kestrel"}` (a bundled invented room; `linnet` is a clean one).
- Limits: 40 documents, 60,000 characters each, 400,000 in all, 2 MB per request. Text only: convert PDFs first (the
document reader block reads scans).
- Categories: `change_of_control`, `anti_assignment`, `exclusivity_noncompete`, `mfn`, `uncapped_liability`,
`debt_covenant`, `ip_assignment_gap`, `undisclosed_litigation`, `customer_concentration`, `key_person` (all by default).
- The JSON response has `flags` (each with `id`, `category`, `severity`, `doc`, `title`, `quote`, `chunk`, `byte_start`,
`byte_end`, `finding`, `why`, `ask`, `grounding`, `status`), `counts`, `not_found`, `markdown` (the memo), `index`
(the index hash, embedder and reranker), `search_receipts`, `record` and `receipts`.
- With `Accept: text/event-stream` (or `"stream": true`) the events are `ready`, `index`, a `receipt` per model call, a
`category` per category, a `grounding` per flag, then `memo`, `report`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
- `GET /madd/info`, `GET /madd/samples` and `GET /attest/signing-key` need no token.
## Example (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
docs = [{"id": p.stem, "title": p.name, "text": p.read_text()} for p in sorted(pathlib.Path("invented_room").glob("*.txt"))]
r = httpx.post(f"{API}/madd/review", json={"room": {"target": "Example Widgets, Inc.", "documents": docs}, "synthetic": True},
headers=H, timeout=900)
r.raise_for_status()
js = r.json()
for f in js["flags"]:
print(f"{f['severity']:>6} {f['label']:<28} {f['doc']} bytes {f['byte_start']}-{f['byte_end']}: \"{f['quote'][:80]}\"")
pathlib.Path("red-flag-memo.md").write_text(js["markdown"])
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa M&A due-diligence red flags: run it yourself (containers)
You are setting up Decosa's due-diligence red-flag review on this machine, so an NDA-bound data room never leaves it. For
ten red-flag categories it searches the room, has an open model name the flags in the room's exact words, checks every
quote word for word against the byte span it cites, checks each finding with a grounding judge, and writes a memo for
counsel and a signed record. Nothing is sent to Decosa's hosted API. It is a triage aid for counsel, not legal advice.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_MADD_SYNTHETIC_ONLY=0` (so this box accepts real rooms), and bind every port to 127.0.0.1. Never set the
gateway route on a box that holds a real data room.
3. Retrieval: either add the `retrieval` service from the decosa-api source (`services/retrieval`, Qwen3-Embedding-0.6B
and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or
set `DECOSA_MADD_RETRIEVAL=bm25` to run with keyword search only (no extra GPU memory; see the eval for the trade-off).
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
5. Check: `curl -fsS http://127.0.0.1:<PORT>/madd/info` shows `synthetic_only: false` and `retrieval.service_reachable:
true` (or mode `bm25`); `GET /attest/signing-key` shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"ma-dd-redflags"}` and send `{"sample": "kestrel"}` to
`POST /madd/review`. Expect at least ten flags, including the MSA's change-of-control clause (MSA-BRI), the demand
letter in the March minutes (MIN-2026Q1), the 3.42x leverage covenant (FIN-FY2025) and the 41.1% customer
(REV-FY2025), and none in the NDA. Then `{"sample": "linnet"}` (a clean room) should give at most one flag. Send the
first record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
client's data room. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 29 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 37 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (66.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (66.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"ma-dd-redflags"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py ma-dd-redflags
Download the mock-data bundle (16 KB, 13 checks)expected.json
Project Kestrel: an invented buyer (Halvard Capital) acquiring an invented target (Brightwater Analytics). Sixteen documents: a customer MSA, an inbound software licence, a reseller agreement, a supply agreement, a county services agreement, a credit agreement, a contractor agreement, the IP assignment register, the CEO employment agreement, an office lease, an NDA, two sets of board minutes, the disclosure schedule, revenue by customer and a financial summary. Planted: change of control in the MSA and a single-trigger CEO payment, a key-person clause, an anti-assignment clause that reaches mergers, exclusivity binding affiliates, MFN pricing, uncapped data-breach liability, a leverage covenant breach (3.42x against 3.00x), a contractor who kept the IP and has no assignment on file, a demand letter in the March minutes that the disclosure schedule leaves out, and one customer at 41.1% of revenue. The run must find them with verbatim quotes, keep the NDA clean, sign a record that verifies and fails when changed, and then return no flags on Project Linnet, a clean room of near misses.
What the rehearsal checks
- the customer MSA's change-of-control termination right is flagged
- the demand letter in the March minutes, missing from the disclosure schedule, is flagged
- the covenant breach is flagged with the figure
- the 41.1% customer is flagged from the revenue table
- the contractor who kept the IP is flagged
- at least ten flags in all
- nothing is flagged in the plain NDA
- the memo is written for counsel and says it is not legal advice
- every search is tied to one index hash
- the signed record verifies
- the record fails once its flag count is changed
- the clean room comes back with at most one flag
- every model call has a signed receipt
Licence: Synthetic: every company, person, clause and number is invented (scripts/madd_rooms.py, decosa_api/verticals/madd/data/rooms.json). Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa M&A due-diligence red flags: run it yourself (containers)
You are setting up Decosa's due-diligence red-flag review on this machine, so an NDA-bound data room never leaves it. For
ten red-flag categories it searches the room, has an open model name the flags in the room's exact words, checks every
quote word for word against the byte span it cites, checks each finding with a grounding judge, and writes a memo for
counsel and a signed record. Nothing is sent to Decosa's hosted API. It is a triage aid for counsel, not legal advice.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_MADD_SYNTHETIC_ONLY=0` (so this box accepts real rooms), and bind every port to 127.0.0.1. Never set the
gateway route on a box that holds a real data room.
3. Retrieval: either add the `retrieval` service from the decosa-api source (`services/retrieval`, Qwen3-Embedding-0.6B
and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or
set `DECOSA_MADD_RETRIEVAL=bm25` to run with keyword search only (no extra GPU memory; see the eval for the trade-off).
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
5. Check: `curl -fsS http://127.0.0.1:<PORT>/madd/info` shows `synthetic_only: false` and `retrieval.service_reachable:
true` (or mode `bm25`); `GET /attest/signing-key` shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"ma-dd-redflags"}` and send `{"sample": "kestrel"}` to
`POST /madd/review`. Expect at least ten flags, including the MSA's change-of-control clause (MSA-BRI), the demand
letter in the March minutes (MIN-2026Q1), the 3.42x leverage covenant (FIN-FY2025) and the 41.1% customer
(REV-FY2025), and none in the NDA. Then `{"sample": "linnet"}` (a clean room) should give at most one flag. Send the
first record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
client's data room. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierM&A due-diligence red flags on GeForce RTX 5090: use the Lite · BM25 retrieval, no retrieval service tier
The standard tier does not fit: Needs about 37 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (does not fit): Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the 4B reranker (about 9 GB) is too tight; run the lite tier (BM25, no retrieval service) or put the retrieval service on a second card.
Lite · BM25 retrieval, no retrieval service: what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Categories and queries, the quote gate, mergi...: decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17). CPU. Runs on CPU (vram_gb 0 in stack.json).
- One call per category: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for M&A due-diligence red flags, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up M&A due-diligence red flags on my hardware Fetch https://decosa.ai/prompts/ma-dd-redflags-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=ma-dd-redflags) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · BM25 retrieval, no retrieval service (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Categories and queries, the quote gate, mergi...: decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17), CPU - One call per category: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/ma-dd-redflags-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 93 s · ~$0.003 per run · 3 receipts
Loading the nightly status…
Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running retrieval service, local signing; torn down after
Measured cost to run: about $0.019 per data room (hosted, 27 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 13/13 in 101 s (Kestrel with its planted flags, the record verified and failed when changed, the clean Linnet room with no flags; every receipt attested). The retrieval service was not built from the compose file here: the running one was used.
Known limits (5)
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- Measured on five small invented rooms written by the same author as the prompts, and 28 public contracts; not on a real data room or against a lawyer's issues list.
- Ten categories only; tax, employment, data protection, environmental and sanctions issues are not looked for.
- Text only: scanned documents must go through the document reader block first.
- 'No flag found' means nothing turned up in the passages the search returned for that category; a run where model calls failed is marked incomplete.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B plus about 10 GB for the embedder and reranker; the quote checks, memo and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the m&a due-diligence red flags API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
A data room in, a red-flag memo for counsel out: every flag quotes the exact words and the byte span it came from.
For deal counsel, corporate development and private-equity deal teams doing buy-side diligence. Give it a data room as text: contracts, board minutes, the disclosure schedule, revenue by customer, financial summaries. The evidence retrieval block cuts every document into chunks with byte offsets and hashes the corpus into one index hash. For each of ten red-flag categories (change of control, anti-assignment, exclusivity and non-compete, MFN pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration, key person), fixed queries are searched and reranked, and Qwen3.8-27B reads the best passages and names the flags in their exact words. Code checks every quote word for word against the passage it cites and drops what it cannot find; the grounding block judges each finding against that passage. You get a Markdown memo (flag, severity, document, quote, byte span, why it matters, what to ask the seller) and a signed record of hashes. A triage aid for counsel, not legal advice.
- Deployment
- Self-host first
- Regulatory
- Not a regulated activity in itself, but the output feeds legal advice, so a lawyer must review it. Written 27 Sep 2026. ABA Formal Opinion 512, Generative Artificial Intelligence Tools (29 Jul 2024; https://www.americanbar.org/content/dam/aba/administrative/professional_responsibility/ethics-opinions/aba-formal-opinion-512.pdf) reads the duties of competence (Model Rule 1.1), confidentiality (1.6) and supervision (5.1, 5.3) as requiring lawyers to understand a tool's limits, check its output and protect client information put into it (unverified: the ABA site blocked our fetch on 27 Sep 2026, so the date and summary are from the published opinion as we recall it; read it before relying on this). This tool supports those duties by quoting its sources word for word and running on the firm's own box; it does not replace the lawyer's review. Data rooms are usually under NDA: check the NDA before sending any document to a hosted service, which is why the hosted demo takes invented rooms only.
Text description
A data room goes to decosa-api. The evidence retrieval block cuts the documents into chunks with byte offsets, embeds them and hashes the corpus into an index hash. For each of ten categories, fixed queries are searched and reranked, with a signed receipt per search. Qwen3.8-27B reads the best passages and names flags with the exact words. Code checks each quote against the cited bytes and drops any not found; the grounding block judges each finding. The output is a memo for counsel and a signed record of hashes.
At a glance
- Data retention
- Nothing stored: the documents live in memory for the request. The signed record holds hashes (documents, index, quotes) and receipt ids, never document text; logs carry counts only.
- What leaves the box
- Self-hosted on the direct route: nothing. The model, the retrieval service and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented rooms are accepted.
- What it checks
- Ten categories: change of control, anti-assignment, exclusivity and non-compete, MFN pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration, key person.
- What it is not
- Legal advice or a materiality call. A triage aid: every flag quotes the data room so counsel can confirm it in seconds.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
BM25 retrieval, no retrieval service
The same review with DECOSA_MADD_RETRIEVAL=bm25: keyword search picks the passages, no embedder or reranker, so only the language model needs a GPU. Same quote gate, grounding and record.
- Models
- decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x GPU for Qwen3.8-27B
- Quality evidence
- Planted flags caught, held-out rooms, BM25 only20 / 20, 0 in the clean roomdecosa-api docs/evals/ma-dd-redflags.md, test split (direct route), 27 Sep 2026
- CUAD public contracts: anti-assignment / exclusivity found, BM25 only11/12 (precision 0.917) / 4/7decosa-api docs/evals/ma-dd-redflags.md, CUAD check
- Latency
- measured: seconds per data room on the direct route
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted; searches are BM25 in code, so there are no embed or rerank attestations.
- In the hosted demo
Standard
retrieval service + the model (hosted demo)
The evidence retrieval block (Qwen3-Embedding-0.6B + Qwen3-Reranker-4B) picks the passages per category, Qwen3.8-27B names the flags, code checks the quotes and the grounding block checks the findings. This is what the hosted demo runs, on invented rooms only.
- Models
- decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)
- Qwen3.8-27B (NVFP4)
- Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards)
- Quality evidence
- Planted flags caught, held-out rooms (Osprey, Heron)20 / 20decosa-api docs/evals/ma-dd-redflags.md, test split, 27 Sep 2026
- Flags raised in the clean held-out room (Wren)0decosa-api docs/evals/ma-dd-redflags.md, test split
- Quotes found word for word at the cited byte span22 / 22 (held-out memo flags)decosa-api docs/evals/ma-dd-redflags.md, test split
- Planted flags caught with 48 public contracts added as distractors (40-document rooms)20 / 20, 0 flags in the added contractsdecosa-api docs/evals/ma-dd-redflags.md, stress split (direct route)
- CUAD public contracts: anti-assignment / exclusivity / change of control found9/12 (precision 0.75) / 3/7 / 2/2decosa-api docs/evals/ma-dd-redflags.md, CUAD check, 28 contracts
- Latency
- measured: under a minute per data room on the shared gateway route, a few minutes under heavy load; about a minute for a larger room self-hosted.
- Verification
- Proof: strongEvery Qwen3.8 call is a separate gateway call with a gateway-signed receipt; every search has a signed search receipt with the index hash; the record lists them all.
Also runs on
- Scanned data rooms (document reader)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17), Qwen3.8-27B (NVFP4), Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)not builtRead scanned PDFs into text with page and box cites first, then review. The document reader block exists; the two are not wired together here yet. Hardware: 1x RTX PRO 6000 96 GB (estimate).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17) 0 GBProof: partial | LiteStandardAlternate | 0 GB | Proof: partial | |
| ||||
One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each findingQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandardAlternate | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab) 0.6B + 4B · 9 GBProof: partialIn the hosted demo | StandardAlternate | 0.6B + 4B · 9 GB | Proof: partialIn the hosted demo | |
| ||||
Tools, services and hardware
Tools
Public contracts with expert clause labels, eval only: 24 contracts added to each held-out room as distractors, and 28 more scored against CUAD's labels (credit: The Atticus Project).
- ABA Formal Opinion 512, Generative Artificial Intelligence Tools (29 Jul 2024) (opens in a new tab)ABA publication
Lawyers' duties when using AI tools; cited in the regulatory note (not fetched: the site blocked us on 27 Sep 2026).
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /madd/info, /madd/samples; POST /madd/review (SSE or JSON); POST /record/verify. Keeps no document text.
- decosa-retrieval:8499
built from services/retrieval (no published image yet)Embedder and reranker for the evidence retrieval block, on the same GPU box; nothing leaves it.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the eval and the hosted demo ran Qwen3.8-27B through the shared gateway on one card and the retrieval service on another GPU (about 10 GB).
- 1x RTX 5090 32 GB Does not fit
Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the 4B reranker (about 9 GB) is too tight; run the lite tier (BM25, no retrieval service) or put the retrieval service on a second card.
- CPU only Does not fit
The model needs a GPU. The quote checks, memo and record run on CPU.
Latency per lane
- one 16-document room, 10 categories with grounding, hosted gateway route32.1 s
Measuredmeasured on our server 2026-09-27: held-out rooms Osprey 34.5 s, Heron 32.1 s, Wren (clean) 20.8 s, gateway shared with other workloads
- the same, under heavy gateway load218.5 s
Measuredmeasured on our server 2026-09-27: dev room Kestrel, 22 model calls
- smoke test: 3 categories, no grounding, hosted gateway route92.9 s
Measuredmeasured on our server 2026-09-27, three runs 86-93 s while the gateway was slow (one run lost two calls to gateway errors and was marked incomplete)
- 40-document room (about 550 chunks), self-hosted direct route52.0 s
Measuredmeasured on our server 2026-09-27: stress rooms 47 s and 58 s
Notes
- On three held-out invented rooms (20 planted red flags, plus a clean room of near misses) it caught 20 of 20 in the right category and raised no flag in the clean room; every quote was found word for word at its byte span.
- Those rooms are at the ceiling: plain BM25 search and reading the whole room also found all 20. What retrieval buys is bounded prompts (a third of the tokens of reading the room) and a signed trail of which corpus version and passages the model saw.
- On 28 public CUAD contracts scored against expert labels it is weaker: anti-assignment 9/12 found (precision 0.75), exclusivity and non-compete 3/7, change of control 2/2; BM25 did as well or better there (11/12, 4/7).
- A triage aid: ten categories, never says a room is clean, and every flag is for counsel to confirm.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa M&A due-diligence red flags on this machine
You are setting up a self-hosted due-diligence red-flag review on this Linux machine for a deal team or law firm. It
reads a data room as text (contracts, board minutes, the disclosure schedule, revenue by customer, financial summaries)
and returns:
- a red-flag memo for counsel: for ten categories (change of control, anti-assignment, exclusivity and non-compete, MFN
pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule,
customer concentration, key person), each flag with the exact words quoted, the document, chunk and byte span, why it
matters and what to ask the seller;
- a signed, hash-chained record of hashes: every document, the retrieval index hash, every search receipt id, every
flag's byte span and quote hash, and every model receipt id.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Data rooms are under NDA. On this box nothing leaves the machine: the language model, the retrieval models and the
checks all run here. Never point it at a hosted gateway while it holds a real room.
- This is a triage aid for counsel, not legal advice. "No flag found" means nothing turned up in the passages searched,
not that the room is clean.
Repeat both points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `retrieval` | built from `decosa-api/services/retrieval` | `Qwen/Qwen3-Embedding-0.6B` @ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` + `Qwen/Qwen3-Reranker-4B` @ `22e683669bc0f0bd69640a1354a6d0aebcfeede5`, Apache-2.0 | internal 8499 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB (the model about 20 GB plus KV cache, the retrieval
models about 10 GB) and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main` and
`LLM_GPU_UTIL=0.70`; on 48 GB also `LLM_MAX_LEN=32768`. Not measured.
- Under 48 GB: tell me, and use the BM25 option in step 3 (no `retrieval` service).
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
`sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 80 GB of free disk.
## 2. Get the images and the source
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, build from source once `decosa-api` is published: clone it into `~/decosa-dd/decosa-api`, run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from
its compose file. If neither works, stop and tell me. The `retrieval` service is always built from the clone
(`services/retrieval`), so clone it either way.
## 3. Write the compose file
Create `~/decosa-dd/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.72
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example LLP deal team>"
```
Create `~/decosa-dd/docker-compose.yml` with exactly these services:
```yaml
name: decosa-dd
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
retrieval:
build:
context: ./decosa-api/services/retrieval
dockerfile_inline: |
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
COPY *.py .
CMD ["python", "server.py"]
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
restart: unless-stopped
environment:
RETRIEVAL_HOST: 0.0.0.0
RETRIEVAL_PORT: "8499"
RETRIEVAL_EMBED: qwen3-emb-0.6b
RETRIEVAL_RERANK: qwen3-rr-4b
HF_HOME: /root/.cache/huggingface
volumes: [hf-cache:/root/.cache/huggingface]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, retrieval: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_RETRIEVAL_URL: http://retrieval:8499
DECOSA_MADD_SYNTHETIC_ONLY: "0" # this box accepts real data rooms
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a full review uses about 7,500 generated tokens
DECOSA_SESSION_TTL_S: "28800"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
The retrieval service downloads its two models (about 9 GB) into the shared `hf-cache` volume on first start. It uses
the pinned revisions in `services/retrieval/models.py`.
**No room for the retrieval models?** Delete the `retrieval` service and its `depends_on` entry, drop
`DECOSA_RETRIEVAL_URL`, and add `DECOSA_MADD_RETRIEVAL: bm25` to the api. Passages are then picked by keyword search
only; the quote gate, grounding and record are unchanged.
Use the named volume `decosa-data` exactly as written: a root-owned host bind mount makes the API fail on
`/data/keys.sqlite`.
Run `docker compose up -d --build`, then poll `docker compose ps` until every service is healthy (the LLM takes 5-10
minutes the first time). `curl -s localhost:8445/madd/info | jq '{synthetic_only, retrieval}'` should show
`synthetic_only: false` and `service_reachable: true`.
## 4. Smoke test on the bundled invented rooms
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"ma-dd-redflags"}' | jq -r .token)
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"kestrel"}' > /tmp/dd.json
jq -r '.flags[] | "\(.severity)\t\(.category)\t\(.doc)\t\(.quote[0:70])"' /tmp/dd.json
jq '{record}' /tmp/dd.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"linnet"}' | jq '.counts'
```
Pass if Kestrel gives at least ten flags, including a change-of-control flag in `MSA-BRI`, an undisclosed-litigation
flag in `MIN-2026Q1`, a debt-covenant flag quoting `3.42x` and a customer-concentration flag in `REV-FY2025`, none in
`NDA-BRI`; every receipt is `attested`; the record verifies; and the clean Linnet room gives at most one flag.
## 5. Review a real room
Put each document in a text file (convert PDFs first; scanned ones through the document reader), then:
```bash
python3 - <<'PY' > room.json
import json, pathlib
docs = [{"id": p.stem[:60], "title": p.name, "text": p.read_text()} for p in sorted(pathlib.Path("room").glob("*.txt"))]
print(json.dumps({"room": {"target": "Example Widgets, Inc.", "buyer": "Example Holdings", "documents": docs}}))
PY
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @room.json > review.json
jq -r .markdown review.json > red-flag-memo.md
jq .record review.json > red-flag-record.json
```
Limits per run: 40 documents, 60,000 characters each, 400,000 in all. Split a bigger room by folder (contracts, corporate,
finance) and run each.
## 6. Point tools at the local API
- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /madd/review` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /madd/info` lists the
categories with their definitions and near-miss rules.
- Keep the memo and the record with the deal file. Anyone can re-check the record with `POST /record/verify` against the
key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send the room's passages to the hosted Decosa
API. Leave it off for real data rooms.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A data room in, a red-flag memo for counsel out: every flag quotes the exact words and the byte span it came from.
- Who it's for
- Deal counsel, corporate development teams and private-equity deal teams doing buy-side diligence.
- Where it runs
- Self-host for a real data room; the hosted demo takes invented rooms only
- Key numbers
- 9 / 12 CUAD anti-assignment clauses found (held out, n = 12)
- 3 / 7 CUAD exclusivity and non-compete clauses found (held out, n = 7)
- 20 / 20 Planted red flags caught (test split, n = 20)
- 92.9 s Median end-to-end run, hosted (QA sweep 2026-09-27)
- Models
- Qwen3.8-27B (one call per category, grounding of each finding) · Qwen3-Embedding + Qwen3-Reranker (the evidence retrieval block)
- Where
- Self-host for a real data room; the hosted demo takes invented rooms only
- Checks
- Receipt per model call; signed search receipts with the index hash; every quote found word for word at the cited byte span; findings grounded; signed hash-chained record
- Industry
- Legal · Finance and insurance
- Runs
- Self-host
- Output
- Signed record or verdict · Notes, reports and drafts
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Evidence retrieval · Grounding · Signed record
Questions people ask
Does our data room leave our network?
Not when you self-host: the model, the retrieval service and the checks all run on your own GPU box. The hosted demo accepts invented data rooms only and keeps nothing.
Can it make up a clause?
A flag reaches the memo only if its quote is found word for word in the data room; code checks every quote against the passage and byte span it cites and drops what it cannot find. The finding sentence is then checked against that passage by a grounding judge.
How many red flags does it catch?
On two held-out invented data rooms it caught 20 of 20 planted red flags and raised none in a clean room of near misses. On 28 public contracts scored against expert labels it found 9 of 12 anti-assignment clauses and 3 of 7 exclusivity clauses. No real deal has been measured yet.
Which red flags does it look for?
Change of control, anti-assignment, exclusivity and non-compete, most-favoured-customer pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration and key-person clauses.
Is this legal advice?
No. It is a triage aid for counsel: it finds and quotes, and a lawyer decides what matters and what to ask the seller.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about M&A due-diligence red flags
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…