Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Write a due-diligence red-flag memo

A red-flag memo for counsel where every flag quotes the exact words from the data room, with the document, severity and what to ask the seller.

Held-out test20 / 20Planted red flags caught (held-out invented data rooms)
On production93 smedian on production (2026-09-27); slower when the service is busy
List price~$0.019 per data roommeasured, at list price

Built on: Evidence retrieval, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B plus about 10 GB for the embedder and reranker; the quote checks, memo and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the m&a due-diligence red flags API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
ma-dd-redflags

Use the hosted API

# Decosa M&A due-diligence red flags: use the hosted API

You are wiring Decosa's due-diligence red-flag review into this project (a deal-room tool, a diligence tracker or a
script). It takes a data room as text and returns a red-flag memo for counsel: each flag with its category, severity,
the document, the exact words quoted, the chunk and byte span, why it matters and what to ask the seller, plus a signed
record. Every model call has a signed receipt and every search has a signed receipt naming the index hash. Use only what
is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes invented data rooms only** (`"synthetic": true` is required). Real data rooms are under NDA:
  use the self-hosted version for them. Never send a real client document here.
- This is a triage aid for counsel, not legal advice. Never present its memo as a clean bill of health: "no flag found"
  means nothing turned up in the passages searched.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "ma-dd-redflags"}` returns `{"token", "expires_at", "budget"}`.
   The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response). Over a limit you get HTTP 429
   with `Retry-After`. A demo token runs one review at a time (409 otherwise).
3. A full review needs about 7,500 generated tokens left in the budget (402 otherwise); fewer categories need less.

## Review a room
- `POST /madd/review` (token). Body: `{"room": {"target": "Example Widgets, Inc.", "buyer"?: "...", "title"?: "...",
  "documents": [{"id"?, "title"?, "text"}]}, "synthetic": true, "categories"?: [...], "ground"?: true, "stream"?: false}`
  or `{"sample": "kestrel"}` (a bundled invented room; `linnet` is a clean one).
  - Limits: 40 documents, 60,000 characters each, 400,000 in all, 2 MB per request. Text only: convert PDFs first (the
    document reader block reads scans).
  - Categories: `change_of_control`, `anti_assignment`, `exclusivity_noncompete`, `mfn`, `uncapped_liability`,
    `debt_covenant`, `ip_assignment_gap`, `undisclosed_litigation`, `customer_concentration`, `key_person` (all by default).
  - The JSON response has `flags` (each with `id`, `category`, `severity`, `doc`, `title`, `quote`, `chunk`, `byte_start`,
    `byte_end`, `finding`, `why`, `ask`, `grounding`, `status`), `counts`, `not_found`, `markdown` (the memo), `index`
    (the index hash, embedder and reranker), `search_receipts`, `record` and `receipts`.
  - With `Accept: text/event-stream` (or `"stream": true`) the events are `ready`, `index`, a `receipt` per model call, a
    `category` per category, a `grounding` per flag, then `memo`, `report`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
- `GET /madd/info`, `GET /madd/samples` and `GET /attest/signing-key` need no token.

## Example (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
docs = [{"id": p.stem, "title": p.name, "text": p.read_text()} for p in sorted(pathlib.Path("invented_room").glob("*.txt"))]
r = httpx.post(f"{API}/madd/review", json={"room": {"target": "Example Widgets, Inc.", "documents": docs}, "synthetic": True},
               headers=H, timeout=900)
r.raise_for_status()
js = r.json()
for f in js["flags"]:
    print(f"{f['severity']:>6}  {f['label']:<28} {f['doc']}  bytes {f['byte_start']}-{f['byte_end']}: \"{f['quote'][:80]}\"")
pathlib.Path("red-flag-memo.md").write_text(js["markdown"])
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa M&A due-diligence red flags: run it yourself (containers)

You are setting up Decosa's due-diligence red-flag review on this machine, so an NDA-bound data room never leaves it. For
ten red-flag categories it searches the room, has an open model name the flags in the room's exact words, checks every
quote word for word against the byte span it cites, checks each finding with a grounding judge, and writes a memo for
counsel and a signed record. Nothing is sent to Decosa's hosted API. It is a triage aid for counsel, not legal advice.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_MADD_SYNTHETIC_ONLY=0` (so this box accepts real rooms), and bind every port to 127.0.0.1. Never set the
   gateway route on a box that holds a real data room.
3. Retrieval: either add the `retrieval` service from the decosa-api source (`services/retrieval`, Qwen3-Embedding-0.6B
   and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or
   set `DECOSA_MADD_RETRIEVAL=bm25` to run with keyword search only (no extra GPU memory; see the eval for the trade-off).
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
5. Check: `curl -fsS http://127.0.0.1:<PORT>/madd/info` shows `synthetic_only: false` and `retrieval.service_reachable:
   true` (or mode `bm25`); `GET /attest/signing-key` shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"ma-dd-redflags"}` and send `{"sample": "kestrel"}` to
   `POST /madd/review`. Expect at least ten flags, including the MSA's change-of-control clause (MSA-BRI), the demand
   letter in the March minutes (MIN-2026Q1), the 3.42x leverage covenant (FIN-FY2025) and the 41.1% customer
   (REV-FY2025), and none in the NDA. Then `{"sample": "linnet"}` (a clean room) should give at most one flag. Send the
   first record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long each run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
client's data room. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 29 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 37 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (66.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (66.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.

  • Apple M5 Max, 64 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"ma-dd-redflags"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py ma-dd-redflags

Download the mock-data bundle (16 KB, 13 checks)expected.json

Project Kestrel: an invented buyer (Halvard Capital) acquiring an invented target (Brightwater Analytics). Sixteen documents: a customer MSA, an inbound software licence, a reseller agreement, a supply agreement, a county services agreement, a credit agreement, a contractor agreement, the IP assignment register, the CEO employment agreement, an office lease, an NDA, two sets of board minutes, the disclosure schedule, revenue by customer and a financial summary. Planted: change of control in the MSA and a single-trigger CEO payment, a key-person clause, an anti-assignment clause that reaches mergers, exclusivity binding affiliates, MFN pricing, uncapped data-breach liability, a leverage covenant breach (3.42x against 3.00x), a contractor who kept the IP and has no assignment on file, a demand letter in the March minutes that the disclosure schedule leaves out, and one customer at 41.1% of revenue. The run must find them with verbatim quotes, keep the NDA clean, sign a record that verifies and fails when changed, and then return no flags on Project Linnet, a clean room of near misses.

What the rehearsal checks
  • the customer MSA's change-of-control termination right is flagged
  • the demand letter in the March minutes, missing from the disclosure schedule, is flagged
  • the covenant breach is flagged with the figure
  • the 41.1% customer is flagged from the revenue table
  • the contractor who kept the IP is flagged
  • at least ten flags in all
  • nothing is flagged in the plain NDA
  • the memo is written for counsel and says it is not legal advice
  • every search is tied to one index hash
  • the signed record verifies
  • the record fails once its flag count is changed
  • the clean room comes back with at most one flag
  • every model call has a signed receipt

Licence: Synthetic: every company, person, clause and number is invented (scripts/madd_rooms.py, decosa_api/verticals/madd/data/rooms.json). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa M&A due-diligence red flags: run it yourself (containers)

You are setting up Decosa's due-diligence red-flag review on this machine, so an NDA-bound data room never leaves it. For
ten red-flag categories it searches the room, has an open model name the flags in the room's exact words, checks every
quote word for word against the byte span it cites, checks each finding with a grounding judge, and writes a memo for
counsel and a signed record. Nothing is sent to Decosa's hosted API. It is a triage aid for counsel, not legal advice.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_MADD_SYNTHETIC_ONLY=0` (so this box accepts real rooms), and bind every port to 127.0.0.1. Never set the
   gateway route on a box that holds a real data room.
3. Retrieval: either add the `retrieval` service from the decosa-api source (`services/retrieval`, Qwen3-Embedding-0.6B
   and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or
   set `DECOSA_MADD_RETRIEVAL=bm25` to run with keyword search only (no extra GPU memory; see the eval for the trade-off).
4. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
5. Check: `curl -fsS http://127.0.0.1:<PORT>/madd/info` shows `synthetic_only: false` and `retrieval.service_reachable:
   true` (or mode `bm25`); `GET /attest/signing-key` shows this box's public key.
6. Smoke test: get a token with `POST /demo/session {"vertical":"ma-dd-redflags"}` and send `{"sample": "kestrel"}` to
   `POST /madd/review`. Expect at least ten flags, including the MSA's change-of-control clause (MSA-BRI), the demand
   letter in the March minutes (MIN-2026Q1), the 3.42x leverage covenant (FIN-FY2025) and the 41.1% customer
   (REV-FY2025), and none in the NDA. Then `{"sample": "linnet"}` (a clean room) should give at most one flag. Send the
   first record to `POST /record/verify`: `ok` must be true.
7. Report back: the public key, the flag counts, and how long each run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
client's data room. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierM&A due-diligence red flags on GeForce RTX 5090: use the Lite · BM25 retrieval, no retrieval service tier

The standard tier does not fit: Needs about 37 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (does not fit): Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the 4B reranker (about 9 GB) is too tight; run the lite tier (BM25, no retrieval service) or put the retrieval service on a second card.

Lite · BM25 retrieval, no retrieval service: what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Categories and queries, the quote gate, mergi...: decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • One call per category: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for M&A due-diligence red flags, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up M&A due-diligence red flags on my hardware

Fetch https://decosa.ai/prompts/ma-dd-redflags-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=ma-dd-redflags)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · BM25 retrieval, no retrieval service (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Categories and queries, the quote gate, mergi...: decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17), CPU
- One call per category: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/ma-dd-redflags-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 93 s · ~$0.003 per run · 3 receipts

Loading the nightly status…

Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running retrieval service, local signing; torn down after

Measured cost to run: about $0.019 per data room (hosted, 27 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The rehearsal bundle passed 13/13 in 101 s (Kestrel with its planted flags, the record verified and failed when changed, the clean Linnet room with no flags; every receipt attested). The retrieval service was not built from the compose file here: the running one was used.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
  • Measured on five small invented rooms written by the same author as the prompts, and 28 public contracts; not on a real data room or against a lawyer's issues list.
  • Ten categories only; tax, employment, data protection, environmental and sanctions issues are not looked for.
  • Text only: scanned documents must go through the document reader block first.
  • 'No flag found' means nothing turned up in the passages the search returned for that category; a run where model calls failed is marked incomplete.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B plus about 10 GB for the embedder and reranker; the quote checks, memo and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the m&a due-diligence red flags API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

A data room in, a red-flag memo for counsel out: every flag quotes the exact words and the byte span it came from.

For deal counsel, corporate development and private-equity deal teams doing buy-side diligence. Give it a data room as text: contracts, board minutes, the disclosure schedule, revenue by customer, financial summaries. The evidence retrieval block cuts every document into chunks with byte offsets and hashes the corpus into one index hash. For each of ten red-flag categories (change of control, anti-assignment, exclusivity and non-compete, MFN pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration, key person), fixed queries are searched and reranked, and Qwen3.8-27B reads the best passages and names the flags in their exact words. Code checks every quote word for word against the passage it cites and drops what it cannot find; the grounding block judges each finding against that passage. You get a Markdown memo (flag, severity, document, quote, byte span, why it matters, what to ask the seller) and a signed record of hashes. A triage aid for counsel, not legal advice.

Deployment
Self-host first
Regulatory
Not a regulated activity in itself, but the output feeds legal advice, so a lawyer must review it. Written 27 Sep 2026. ABA Formal Opinion 512, Generative Artificial Intelligence Tools (29 Jul 2024; https://www.americanbar.org/content/dam/aba/administrative/professional_responsibility/ethics-opinions/aba-formal-opinion-512.pdf) reads the duties of competence (Model Rule 1.1), confidentiality (1.6) and supervision (5.1, 5.3) as requiring lawyers to understand a tool's limits, check its output and protect client information put into it (unverified: the ABA site blocked our fetch on 27 Sep 2026, so the date and summary are from the published opinion as we recall it; read it before relying on this). This tool supports those duties by quoting its sources word for word and running on the firm's own box; it does not replace the lawyer's review. Data rooms are usually under NDA: check the NDA before sending any document to a hosted service, which is why the hosted demo takes invented rooms only.
Architecture
Text description

A data room goes to decosa-api. The evidence retrieval block cuts the documents into chunks with byte offsets, embeds them and hashes the corpus into an index hash. For each of ten categories, fixed queries are searched and reranked, with a signed receipt per search. Qwen3.8-27B reads the best passages and names flags with the exact words. Code checks each quote against the cited bytes and drops any not found; the grounding block judges each finding. The output is a memo for counsel and a signed record of hashes.

Architecture

At a glance

Data retention
Nothing stored: the documents live in memory for the request. The signed record holds hashes (documents, index, quotes) and receipt ids, never document text; logs carry counts only.
What leaves the box
Self-hosted on the direct route: nothing. The model, the retrieval service and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented rooms are accepted.
What it checks
Ten categories: change of control, anti-assignment, exclusivity and non-compete, MFN pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration, key person.
What it is not
Legal advice or a materiality call. A triage aid: every flag quotes the data room so counsel can confirm it in seconds.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    BM25 retrieval, no retrieval service

    The same review with DECOSA_MADD_RETRIEVAL=bm25: keyword search picks the passages, no embedder or reranker, so only the language model needs a GPU. Same quote gate, grounding and record.

    Models
    • decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x GPU for Qwen3.8-27B
    Quality evidence
    • Planted flags caught, held-out rooms, BM25 only20 / 20, 0 in the clean roomdecosa-api docs/evals/ma-dd-redflags.md, test split (direct route), 27 Sep 2026
    • CUAD public contracts: anti-assignment / exclusivity found, BM25 only11/12 (precision 0.917) / 4/7decosa-api docs/evals/ma-dd-redflags.md, CUAD check
    Latency
    measured: seconds per data room on the direct route
    Verification
    Proof: strongSelf-host onlyEvery model call is receipted; searches are BM25 in code, so there are no embed or rerank attestations.
  • In the hosted demo

    Standard

    retrieval service + the model (hosted demo)

    The evidence retrieval block (Qwen3-Embedding-0.6B + Qwen3-Reranker-4B) picks the passages per category, Qwen3.8-27B names the flags, code checks the quotes and the grounding block checks the findings. This is what the hosted demo runs, on invented rooms only.

    Models
    • decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)
    • Qwen3.8-27B (NVFP4)
    • Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
    Hardware
    1x RTX PRO 6000 96 GB (measured on shared cards)
    Quality evidence
    • Planted flags caught, held-out rooms (Osprey, Heron)20 / 20decosa-api docs/evals/ma-dd-redflags.md, test split, 27 Sep 2026
    • Flags raised in the clean held-out room (Wren)0decosa-api docs/evals/ma-dd-redflags.md, test split
    • Quotes found word for word at the cited byte span22 / 22 (held-out memo flags)decosa-api docs/evals/ma-dd-redflags.md, test split
    • Planted flags caught with 48 public contracts added as distractors (40-document rooms)20 / 20, 0 flags in the added contractsdecosa-api docs/evals/ma-dd-redflags.md, stress split (direct route)
    • CUAD public contracts: anti-assignment / exclusivity / change of control found9/12 (precision 0.75) / 3/7 / 2/2decosa-api docs/evals/ma-dd-redflags.md, CUAD check, 28 contracts
    Latency
    measured: under a minute per data room on the shared gateway route, a few minutes under heavy load; about a minute for a larger room self-hosted.
    Verification
    Proof: strongEvery Qwen3.8 call is a separate gateway call with a gateway-signed receipt; every search has a signed search receipt with the index hash; the record lists them all.

Also runs on

  • Scanned data rooms (document reader)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17), Qwen3.8-27B (NVFP4), Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)not builtRead scanned PDFs into text with page and box cites first, then review. The document reader block exists; the two are not wired together here yet. Hardware: 1x RTX PRO 6000 96 GB (estimate).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)
0 GBProof: partial
One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each findingQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab)
0.6B + 4B · 9 GBProof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /madd/info, /madd/samples; POST /madd/review (SSE or JSON); POST /record/verify. Keeps no document text.

  • decosa-retrieval:8499
    built from services/retrieval (no published image yet)

    Embedder and reranker for the evidence retrieval block, on the same GPU box; nothing leaves it.

  • vLLM (model):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the eval and the hosted demo ran Qwen3.8-27B through the shared gateway on one card and the retrieval service on another GPU (about 10 GB).

  • 1x RTX 5090 32 GB Does not fit

    Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the 4B reranker (about 9 GB) is too tight; run the lite tier (BM25, no retrieval service) or put the retrieval service on a second card.

  • CPU only Does not fit

    The model needs a GPU. The quote checks, memo and record run on CPU.

Latency per lane

  • one 16-document room, 10 categories with grounding, hosted gateway route32.1 s

    Measuredmeasured on our server 2026-09-27: held-out rooms Osprey 34.5 s, Heron 32.1 s, Wren (clean) 20.8 s, gateway shared with other workloads

  • the same, under heavy gateway load218.5 s

    Measuredmeasured on our server 2026-09-27: dev room Kestrel, 22 model calls

  • smoke test: 3 categories, no grounding, hosted gateway route92.9 s

    Measuredmeasured on our server 2026-09-27, three runs 86-93 s while the gateway was slow (one run lost two calls to gateway errors and was marked incomplete)

  • 40-document room (about 550 chunks), self-hosted direct route52.0 s

    Measuredmeasured on our server 2026-09-27: stress rooms 47 s and 58 s

Notes

  • On three held-out invented rooms (20 planted red flags, plus a clean room of near misses) it caught 20 of 20 in the right category and raised no flag in the clean room; every quote was found word for word at its byte span.
  • Those rooms are at the ceiling: plain BM25 search and reading the whole room also found all 20. What retrieval buys is bounded prompts (a third of the tokens of reading the room) and a signed trail of which corpus version and passages the model saw.
  • On 28 public CUAD contracts scored against expert labels it is weaker: anti-assignment 9/12 found (precision 0.75), exclusivity and non-compete 3/7, change of control 2/2; BM25 did as well or better there (11/12, 4/7).
  • A triage aid: ten categories, never says a room is clean, and every flag is for counsel to confirm.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

ma-dd-redflags/assemble-prompt.md213 lines
# Assemble Decosa M&A due-diligence red flags on this machine

You are setting up a self-hosted due-diligence red-flag review on this Linux machine for a deal team or law firm. It
reads a data room as text (contracts, board minutes, the disclosure schedule, revenue by customer, financial summaries)
and returns:
- a red-flag memo for counsel: for ten categories (change of control, anti-assignment, exclusivity and non-compete, MFN
  pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule,
  customer concentration, key person), each flag with the exact words quoted, the document, chunk and byte span, why it
  matters and what to ask the seller;
- a signed, hash-chained record of hashes: every document, the retrieval index hash, every search receipt id, every
  flag's byte span and quote hash, and every model receipt id.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Data rooms are under NDA. On this box nothing leaves the machine: the language model, the retrieval models and the
  checks all run here. Never point it at a hosted gateway while it holds a real room.
- This is a triage aid for counsel, not legal advice. "No flag found" means nothing turned up in the passages searched,
  not that the room is clean.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/ma-dd-redflags.zip (16 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py ma-dd-redflags` (the api image carries the same bundle under /app/rehearsal/ma-dd-redflags/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py ma-dd-redflags --bundle ma-dd-redflags.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the customer MSA's change-of-control termination right is flagged", "the demand letter in the March minutes, missing from the disclosure schedule, is flagged", "the covenant breach is flagged with the figure"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `retrieval` | built from `decosa-api/services/retrieval` | `Qwen/Qwen3-Embedding-0.6B` @ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` + `Qwen/Qwen3-Reranker-4B` @ `22e683669bc0f0bd69640a1354a6d0aebcfeede5`, Apache-2.0 | internal 8499 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB (the model about 20 GB plus KV cache, the retrieval
   models about 10 GB) and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main` and
     `LLM_GPU_UTIL=0.70`; on 48 GB also `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: tell me, and use the BM25 option in step 3 (no `retrieval` service).
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
   NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 80 GB of free disk.

## 2. Get the images and the source

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, build from source once `decosa-api` is published: clone it into `~/decosa-dd/decosa-api`, run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from
its compose file. If neither works, stop and tell me. The `retrieval` service is always built from the clone
(`services/retrieval`), so clone it either way.

## 3. Write the compose file

Create `~/decosa-dd/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.72
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example LLP deal team>"
```

Create `~/decosa-dd/docker-compose.yml` with exactly these services:

```yaml
name: decosa-dd
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  retrieval:
    build:
      context: ./decosa-api/services/retrieval
      dockerfile_inline: |
        FROM python:3.12-slim
        WORKDIR /app
        COPY requirements.txt .
        RUN pip install --no-cache-dir -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
        COPY *.py .
        CMD ["python", "server.py"]
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    restart: unless-stopped
    environment:
      RETRIEVAL_HOST: 0.0.0.0
      RETRIEVAL_PORT: "8499"
      RETRIEVAL_EMBED: qwen3-emb-0.6b
      RETRIEVAL_RERANK: qwen3-rr-4b
      HF_HOME: /root/.cache/huggingface
    volumes: [hf-cache:/root/.cache/huggingface]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, retrieval: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_RETRIEVAL_URL: http://retrieval:8499
      DECOSA_MADD_SYNTHETIC_ONLY: "0"           # this box accepts real data rooms
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "400000"        # per session; a full review uses about 7,500 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

The retrieval service downloads its two models (about 9 GB) into the shared `hf-cache` volume on first start. It uses
the pinned revisions in `services/retrieval/models.py`.

**No room for the retrieval models?** Delete the `retrieval` service and its `depends_on` entry, drop
`DECOSA_RETRIEVAL_URL`, and add `DECOSA_MADD_RETRIEVAL: bm25` to the api. Passages are then picked by keyword search
only; the quote gate, grounding and record are unchanged.

Use the named volume `decosa-data` exactly as written: a root-owned host bind mount makes the API fail on
`/data/keys.sqlite`.

Run `docker compose up -d --build`, then poll `docker compose ps` until every service is healthy (the LLM takes 5-10
minutes the first time). `curl -s localhost:8445/madd/info | jq '{synthetic_only, retrieval}'` should show
`synthetic_only: false` and `service_reachable: true`.

## 4. Smoke test on the bundled invented rooms

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"ma-dd-redflags"}' | jq -r .token)
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"kestrel"}' > /tmp/dd.json
jq -r '.flags[] | "\(.severity)\t\(.category)\t\(.doc)\t\(.quote[0:70])"' /tmp/dd.json
jq '{record}' /tmp/dd.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"linnet"}' | jq '.counts'
```

Pass if Kestrel gives at least ten flags, including a change-of-control flag in `MSA-BRI`, an undisclosed-litigation
flag in `MIN-2026Q1`, a debt-covenant flag quoting `3.42x` and a customer-concentration flag in `REV-FY2025`, none in
`NDA-BRI`; every receipt is `attested`; the record verifies; and the clean Linnet room gives at most one flag.

## 5. Review a real room

Put each document in a text file (convert PDFs first; scanned ones through the document reader), then:

```bash
python3 - <<'PY' > room.json
import json, pathlib
docs = [{"id": p.stem[:60], "title": p.name, "text": p.read_text()} for p in sorted(pathlib.Path("room").glob("*.txt"))]
print(json.dumps({"room": {"target": "Example Widgets, Inc.", "buyer": "Example Holdings", "documents": docs}}))
PY
curl -s $API/madd/review -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @room.json > review.json
jq -r .markdown review.json > red-flag-memo.md
jq .record review.json > red-flag-record.json
```

Limits per run: 40 documents, 60,000 characters each, 400,000 in all. Split a bigger room by folder (contracts, corporate,
finance) and run each.

## 6. Point tools at the local API

- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /madd/review` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /madd/info` lists the
  categories with their definitions and near-miss rules.
- Keep the memo and the record with the deal file. Anyone can re-check the record with `POST /record/verify` against the
  key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send the room's passages to the hosted Decosa
API. Leave it off for real data rooms.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A data room in, a red-flag memo for counsel out: every flag quotes the exact words and the byte span it came from.
Who it's for
Deal counsel, corporate development teams and private-equity deal teams doing buy-side diligence.
Where it runs
Self-host for a real data room; the hosted demo takes invented rooms only
Key numbers
  • 9 / 12 CUAD anti-assignment clauses found (held out, n = 12)
  • 3 / 7 CUAD exclusivity and non-compete clauses found (held out, n = 7)
  • 20 / 20 Planted red flags caught (test split, n = 20)
  • 92.9 s Median end-to-end run, hosted (QA sweep 2026-09-27)
All results, datasets and caveats
Models
Qwen3.8-27B (one call per category, grounding of each finding) · Qwen3-Embedding + Qwen3-Reranker (the evidence retrieval block)
Where
Self-host for a real data room; the hosted demo takes invented rooms only
Checks
Receipt per model call; signed search receipts with the index hash; every quote found word for word at the cited byte span; findings grounded; signed hash-chained record
Output
Signed record or verdict · Notes, reports and drafts
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Does our data room leave our network?

Not when you self-host: the model, the retrieval service and the checks all run on your own GPU box. The hosted demo accepts invented data rooms only and keeps nothing.

Can it make up a clause?

A flag reaches the memo only if its quote is found word for word in the data room; code checks every quote against the passage and byte span it cites and drops what it cannot find. The finding sentence is then checked against that passage by a grounding judge.

How many red flags does it catch?

On two held-out invented data rooms it caught 20 of 20 planted red flags and raised none in a clean room of near misses. On 28 public contracts scored against expert labels it found 9 of 12 anti-assignment clauses and 3 of 7 exclusivity clauses. No real deal has been measured yet.

Which red flags does it look for?

Change of control, anti-assignment, exclusivity and non-compete, most-favoured-customer pricing, uncapped liability, debt covenants, IP assignment gaps, litigation missing from the disclosure schedule, customer concentration and key-person clauses.

Is this legal advice?

No. It is a triage aid for counsel: it finds and quotes, and a lawyer decides what matters and what to ask the seller.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about M&A due-diligence red flags

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.