Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Record F&I add-on disclosures

A signed record per deal of what was said about each add-on, quoted with its time, and every place the deal jacket differs from the conversation.

Held-out test31/32Checks answered correctly on held-out deals (first run)
On production3.3 smedian on production (2026-09-25); slower when the service is busy
List price~$0.27 per 100 conversationsmeasured, at list price

Built on: Speaker diarization, Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the jacket rules and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the auto f&i disclosure record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
fi-disclosure-record

Use the hosted API

# Decosa F&I disclosure record: use the hosted API

You are wiring Decosa's F&I disclosure check into this project. It takes the transcript of a car dealer's F&I-office
conversation and the deal jacket. For each add-on it gives typed answers: was the add-on said to be optional, and was it
presented as required. Each answer is `yes`, `no`, `unclear` or `na`, with a quote and its time. It also checks recording
consent, the total of payments, the lower-payment warning and the 3-day cancel right on used cars. It lists where the
jacket differs from what was said, and reads the written disclosures and banned add-ons of California's CARS Act
(SB 766, in force 1 Oct 2026) from the jacket. It returns a signed, hash-chained record to keep for two years. Every
model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic material only.** Real deal jackets and F&I recordings belong on a self-hosted box (see
  the self-host prompt). Say so wherever this is wired in.
- **Consent.** California requires every party's consent before a confidential conversation is recorded (Penal Code 632).
  Collect it in the UI. Send it with every check; the API refuses a check without it.
- This is triage, not a compliance verdict. Never label a deal "compliant".

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "fi-disclosure-record"}` returns `{"token", "expires_at", "budget"}`.
   The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response), which is about six checks.
   Over a limit you get HTTP 429 with `Retry-After`. A demo token runs one check at a time (409 otherwise).
3. A check needs about 900 generated tokens plus 170 per typed question (402 otherwise).

## Endpoints
- `POST /fi/check` (token). Body: `{"transcript": "[00:05] FINANCE MANAGER: ...", "jacket": {...}, "consent": {"all_parties": true, "method": "verbal_on_recording"}, "speakers"?: {"S01": "Finance manager"}, "title"?: "..."}`.
  - Give exactly one of:
    - `transcript`: text, one utterance per line starting with a time like `[01:23]` and a speaker; WebVTT and SRT also work.
    - `segments`: `[{start, end, speaker, text}]` in seconds, as a diarizer returns them.
    - `sample_id`: from `/fi/samples`. A sample brings its own jacket and consent.
  - `jacket`: `{deal_id?, date?: "YYYY-MM-DD", vehicle: {used: true|false, price, fuel?: gas|diesel|hybrid|plug_in_hybrid|electric, description?}, sale?: {auction, lease_buyout, fleet_or_commercial, gvwr_10k_plus, prior_damage}, financing?: {monthly_payment, term_months, down_payment}, addons: [{id?, name, kind, price, includes_oil_changes?, catalytic_marking?, voids_paint_warranty?, nitrogen_purity_pct?}], documents?: {addon_not_required_written, total_of_payments_written, cancel_form_given, contract_first_page_notice}}`.
    - The add-on `kind` is one of `service_contract`, `gap`, `maintenance`, `tire_wheel`, `nitrogen`, `theft_marking`,
      `surface_protection`, `key_replacement`, `appearance` or `other`.
    - At most 8 add-ons. Leave a document out if you do not know whether it was given.
  - `consent.method`: `verbal_on_recording`, `written`, `signage_and_verbal` or `other`. `all_parties` must be `true`.
  - Limits: 800 transcript lines, 80,000 characters, 512 KB of JSON.
  - JSON response by default:
    - `checklist`: `[{id, rule, title, basis: "conversation"|"jacket", answer, status: "ok"|"flag"|"review"|"na", reason, line?, at?, t_ms?, quote?, evidence?: "quote"|"context", probability?, receipt_ids?}]`
    - `mismatches`: `[{id, kind, status, what, addon?, at?, quote?}]`
    - also `counts`, `cancel: {applies, why, restocking_fee_max?}`, `retain_until`, `record`, `record_check`, `receipts`, `law`, `note`
  - Check ids:
    - `consent_on_record`;
    - per add-on, `addon_optional:<id>` and `addon_required:<id>`;
    - `payment_total_said` and `lower_payment_warning`;
    - on used cars at $50,000 or less, `cancel_right_said` and `cancel_right_misstated`;
    - from the jacket, `doc:<document>` and `no_benefit:<id>:<rule>`.
  - Mismatch kinds: `never_discussed`, `declined_on_jacket`, `no_clear_yes`, `price`, `accepted_not_on_jacket`,
    `monthly_payment`, `term`, `vehicle_price` and `restocking_fee`.
  - With `Accept: text/event-stream` (or `"stream": true`), the events are:
    - `ready`;
    - a `receipt` for each model call, and a `check` for each answer as it completes;
    - then `mismatches`, `jacket`, `result`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, checks, first_bad}`.
- `GET /fi/info`, `GET /fi/samples` and `GET /attest/signing-key` need no token. `POST /fi/transcribe` is self-host only
  (503 here).

## Example: check a deal and print what needs a person's eyes (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"transcript": open("fi-transcript.txt").read(), "jacket": json.load(open("jacket.json")),
        "consent": {"all_parties": True, "method": "verbal_on_recording"}}
r = httpx.post(f"{API}/fi/check", json=body, headers=H, timeout=300)
r.raise_for_status()
js = r.json()
for x in js["checklist"]:
    if x["status"] in ("flag", "review"):
        where = f" ({x['evidence']} at {x['at']}: \"{x.get('quote', '')[:100]}\")" if x.get("at") else ""
        print(f"{x['status'].upper()}: {x['title']}? {x['answer']}{where}")
for m in js["mismatches"]:
    print(f"{m['status'].upper()}: {m['what']}" + (f" (at {m['at']})" if m.get("at") else ""))
json.dump(js["record"], open(f"fi-record-{js['retain_until']}.json", "w"))   # keep until retain_until
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa F&I disclosure record: run it yourself (containers)

You are setting up the Decosa F&I disclosure record on this machine, so deal jackets and F&I recordings never leave it.
It checks each F&I-office conversation against its deal jacket under California's CARS Act (SB 766), and gives typed
answers with quotes and times. It lists jacket mismatches and seals a signed record to keep for two years. Nothing is sent
to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/fi-disclosure-record.zip (3 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py fi-disclosure-record` (the api image carries the same bundle under /app/rehearsal/fi-disclosure-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py fi-disclosure-record --bundle fi-disclosure-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least five conversation checks ran", "every conversation check got a typed answer", "the planted cooling-off denial is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.

   For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
   `DECOSA_FI_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/fi/info` lists the rules with their Civil Code sections.
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"fi-disclosure-record"}`, then
   `POST /fi/check {"sample_id": "tb4-no-consent-no-cooling-off"}`. Expect:
   - `cancel_right_misstated` as `flag`, with a time;
   - a `never_discussed` mismatch for the paint and fabric package;
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.

Recording needs every party's consent (Penal Code 632). This is triage for a person to review, not a compliance verdict
and not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds deal
jackets or F&I recordings. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/fi-disclosure-record-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (61.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"fi-disclosure-record"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py fi-disclosure-record

Download the mock-data bundle (3 KB, 9 checks)expected.json

Eight diarized lines of a role-played used-car F&I conversation (two add-ons) and its deal jacket. The finance manager wrongly says there is no cooling-off period, and one add-on on the jacket is never discussed. Both must be flagged, and the signed record must verify with a two-year retention date.

What the rehearsal checks
  • at least five conversation checks ran
  • every conversation check got a typed answer
  • the planted cooling-off denial is flagged
  • the cooling-off flag points at a time in the recording
  • the add-on never discussed (paint and fabric protection) is flagged
  • the signed record verifies
  • the record carries a retention date
  • a record with its flag count changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic role-play written from California SB 766: Larkspur Point Motors and every person are fictional. Audio spoken by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-fi-demo; the segments are the diarizer output with its errors kept. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa F&I disclosure record: run it yourself (containers)

You are setting up the Decosa F&I disclosure record on this machine, so deal jackets and F&I recordings never leave it.
It checks each F&I-office conversation against its deal jacket under California's CARS Act (SB 766), and gives typed
answers with quotes and times. It lists jacket mismatches and seals a signed record to keep for two years. Nothing is sent
to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/fi-disclosure-record.zip (3 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py fi-disclosure-record` (the api image carries the same bundle under /app/rehearsal/fi-disclosure-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py fi-disclosure-record --bundle fi-disclosure-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least five conversation checks ran", "every conversation check got a typed answer", "the planted cooling-off denial is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
   - set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
   - keep its data on a named volume;
   - bind every port to 127.0.0.1.

   For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
   `DECOSA_FI_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
   downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/fi/info` lists the rules with their Civil Code sections.
   `GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"fi-disclosure-record"}`, then
   `POST /fi/check {"sample_id": "tb4-no-consent-no-cooling-off"}`. Expect:
   - `cancel_right_misstated` as `flag`, with a time;
   - a `never_discussed` mismatch for the paint and fabric package;
   - every receipt `attested`.

   Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.

Recording needs every party's consent (Penal Code 632). This is triage for a person to review, not a compliance verdict
and not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds deal
jackets or F&I recordings. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/fi-disclosure-record-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsAuto F&I disclosure record on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Extraction: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Auto F&I disclosure record, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Auto F&I disclosure record on my hardware

Fetch https://decosa.ai/prompts/fi-disclosure-record-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=fi-disclosure-record)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Extraction: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/fi-disclosure-record-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Auto F&I disclosure record: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Auto F&I disclosure record on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/fi-disclosure-record.zip (3 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py fi-disclosure-record` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least five conversation checks ran", "every conversation check got a typed answer", "the planted cooling-off denial is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Extraction (what was said about each add-on, prices, payment) and one typed yes/no/unclear check per question | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it) | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py fi-disclosure-record`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Built from the same parts as the measured tools; merged after the Mac run, so not run on the Mac yet.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 3.3 s · ~$0.003 per run · 10 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.27 per 100 conversations (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Both smoke tests in the assembly prompt passed against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb4 flagged the cooling-off denial and the add-on never discussed, 10 attested receipts, record verified, 3.8 s; the pasted transcript flagged the service contract as presented as required at 00:05; no consent gave 400. Model-server startup and the diarizer path were not re-run.

Known limits (6)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route), driven from the branch site in headless Chromium, including at 390 px. The production API gets this vertical when the branch merges.
  • Measured on 17 synthetic role-plays written by the building agent, with clean TTS audio. It has not been measured on real F&I recordings (long sessions, crosstalk, Spanish) or with a dealer compliance officer's labels.
  • The Act's add-on and payment disclosures must be in writing. The tool reads them from the jacket as the dealer states them, and cannot judge whether they were clear, conspicuous or in the right language.
  • One systematic miss: a hint at cancelling ('if you did cancel there's a restocking fee') is read as explaining the 3-day right. The misstatement check still flags that call.
  • Audio intake (/fi/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.
  • It does not check advertising or the first written price (1784.41(a)), GAP loan-to-value (1784.42(a)(3)), or add-on payment timing (1784.42(b)).

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the jacket rules and the record need no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the auto f&i disclosure record API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

What the F&I manager said about each add-on, quoted with its time, checked against the deal jacket under California's CARS Act.

For dealer-group compliance officers and CFOs in California. Give it the F&I-office conversation (audio on your own box, or a transcript) and the deal jacket. For each add-on it answers two questions, each yes, no or unclear with a quote and its time: was it said to be optional, and was it presented as required. It also checks recording consent, the total of payments, the lower-payment warning and, on used cars at $50,000 or less, the 3-day right to cancel. It flags add-ons on the contract that were declined or never discussed, prices, payments and terms that differ from what was said, and a restocking fee over the cap. It reads the Act's written disclosures and banned add-ons from the jacket, then seals everything in a signed, hash-chained record to keep for two years. It is triage for a person to review, not a compliance verdict.

Deployment
Self-host first
Regulatory
Checked 25 Sep 2026 against the chaptered text. California SB 766, the California Combating Auto Retail Scams (CARS) Act (Stats. 2025, ch. 354; approved and filed 6 Oct 2025), adds Civil Code 1784.20-1784.44, operative 1 Oct 2026. It bans misrepresenting costs, terms or any aspect of an add-on (1784.40). It requires, in writing, that an add-on is not required and the car can be bought without it, the total paid at a quoted monthly payment, and a warning that lower payments often increase the total (1784.41(b)-(d)). It bans charging for add-ons the buyer cannot benefit from, with seven examples such as oil changes on an EV (1784.42(a)). It gives a 3-day right to cancel on used vehicles sold or leased at $50,000 or less, with a separate disclosure and a restocking fee of 1.5% of the price ($200 minimum, $600 maximum) plus $1 a mile over 250 (up to $150) (1784.31(g), 1784.43). It requires dealers to keep records showing compliance for two years (1784.44). The written disclosures cannot be met by speech: this tool reads them from the jacket and checks whether what was said agrees. Recording an F&I conversation needs the consent of every party (Penal Code 632(a)). The federal FTC CARS Rule was vacated by the Fifth Circuit on 27 Jan 2025 (NADA v. FTC, No. 24-60013) and is context only. Not legal advice.
Architecture
Text description

An F&I-office recording (self-host only) goes through MOSS-Transcribe-Diarize to a timed, speaker-labelled transcript; a typed, WebVTT or SRT transcript can be given instead. With it come the deal jacket (vehicle, price, add-ons, financing, the written disclosures given) and a recording-consent statement. Qwen3.8-27B extracts what was said about each add-on and the payment, and answers one typed question per check (said to be optional? presented as required? consent, total of payments, lower-payment warning, 3-day cancel right), each with a quote that must be found in the transcript. Code compares what was said with the jacket and applies the jacket-only rules (written disclosures, banned add-ons, cancel eligibility, restocking-fee cap). Outputs: the checklist, the mismatches and a signed hash-chained record with its retention date. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it checks
What was said in the F&I office about each add-on, consent, payments and the 3-day cancel right, against the deal jacket. It also checks the jacket's written disclosures and banned add-ons under California SB 766.
What it does not do
It gives no compliance verdict. It does not check advertising or the first written price quote (1784.41(a)), whether a GAP loan-to-value is lawful, add-on payment timing, or whether a written disclosure was clear, conspicuous and in the negotiation language.
Consent
Every run needs a statement that all parties agreed to the recording (Penal Code 632). The statement goes into the signed record, and a check looks for the consent on the recording itself.
Data kept
Nothing on the server. Transcripts, jackets and records live in memory for the request, and logs carry counts and timings only. You keep the signed record for two years (1784.44(a)).
What leaves the box (hosted demo)
The transcript and the jacket's add-on list go to Qwen3.8-27B through our gateway. The gateway's receipts hold hashes, not text. Self-hosted, nothing leaves.
Model calls per deal
One extraction, plus one typed check per question: 2 per add-on, 3 general, and 2 more on used cars at $50,000 or less. That is 10 for a used car with two add-ons.
Typical run cost
A fraction of a cent at the gateway list price for a used-car deal with add-ons (measured on the smoke sample). Each run shows its own measured cost.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same checks on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • typed checks correct / planted problems foundnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B extracts and answers every typed check; MOSS-Transcribe-Diarize turns recordings into transcripts. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    • MOSS-Transcribe-Diarize 0.9B
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • held-out test B, 4 synthetic conversations: typed checks correct / planted problems found / mismatches found31/32 / 8/8 / 3/3decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25, gateway route, run 1; data and prompts written by the building agent, prompts frozen on a separate 3-conversation dev split
    • test set, 10 conversations: typed checks correct (run 1 before two fixes / after)76/78 / 77/78; planted 12/13; mismatches 6/6decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25, gateway route
    • false alarms on clean conversations (flag or review)test 0/30 after fixes (2/30 before); test B 1/14decosa-api docs/evals/fi-disclosure-record.md, measured on our server 2026-09-25
    • 5 conversations on real speech-recognition transcripts of synthetic audio: checks / planted / cited time within 3 s45/45 / 10/10 / 21/21 (repeat 45/45, 10/10, 20/21)decosa-api docs/evals/fi-disclosure-record.md (asr runs), measured on our server 2026-09-26 on audio re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: repeat 44/45, 9/10, 19/21), gateway route
    • real F&I recordings reviewed by a dealer compliance officernot measured yet
    Latency
    measured on our server under a shared gateway: seconds per text conversation, a little longer per audio transcript
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger model for long, many-speaker F&I sessions.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • typed checks correct / planted problems foundnot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Call recordings stay on your own hardware, never on community providers. Not served yet.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • typed checks correct / planted problems foundnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Extraction (what was said about each add-on, prices, payment) and one typed yes/no/unclear check per questionQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialSelf-host only
Lite tier: the same extraction and checks on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger model for long, many-speaker F&I sessionsDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Transcript and jacket intake, the checks, the jacket rules, signing and the HTTP API (/fi/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

  • decosa-diarize:8092

    Optional, self-host only: MOSS-Transcribe-Diarize for transcripts made from recordings (DECOSA_FI_AUDIO=1). No published image yet; built from services/diarize.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server; the diarizer runs on the other card there.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    Transcript parsing, finalize (diffs, provenance, the signed record) and verification need no GPU; the check itself needs the model.

Latency per lane

  • check of a 10-16 line conversation with 2 add-ons, busy shared gateway3.7 s

    Measureddecosa-api docs/evals/fi-disclosure-record.md, test runs on our server 2026-09-25 (median 3.5-4.2 s per conversation), gateway route

  • check of a diarized audio transcript, busy shared gateway7.5 s

    Measureddecosa-api docs/evals/fi-disclosure-record.md, asr runs on our server 2026-09-25 (medians 7.5 s and 16.6 s)

  • transcribe a 30-84 s recording (self-host)4.0 s

    Measuredmeasured on our server 2026-09-25, MOSS-Transcribe-Diarize on an RTX PRO 6000 (3-7 s)

  • jacket rules and the signed record50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

fi-disclosure-record/assemble-prompt.md163 lines
# Assemble the Decosa F&I disclosure record on this machine

You are setting up a self-hosted F&I disclosure checker on this Linux machine for a car dealer. It takes the transcript of an F&I-office conversation and the deal jacket. For each add-on it gives typed answers (yes, no or unclear, with a quote and a time): was the add-on said to be optional, and was it presented as required. It flags add-ons on the contract that were declined or never discussed, and prices, payments or terms that differ from what was said. It reads the written disclosures and banned add-ons of California's CARS Act (SB 766, Civil Code 1784.20-1784.44, in force 1 Oct 2026) from the jacket, and seals everything in a signed, hash-chained record to keep for two years (1784.44(a)). With the optional diarizer it turns a recording into a timed, speaker-labelled transcript. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Deal jackets hold credit applications and personal data. Keep everything on this machine: the model route stays local (`direct`), and nothing goes to a hosted service.
- California requires every party's consent before a confidential conversation is recorded (Penal Code 632). Each check needs a consent statement, and the statement goes into the record.
- This is triage for a person to review. It is not a compliance verdict and not legal advice.

Repeat these points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/fi-disclosure-record.zip (3 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py fi-disclosure-record` (the api image carries the same bundle under /app/rehearsal/fi-disclosure-record/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py fi-disclosure-record --bundle fi-disclosure-record.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least five conversation checks ran", "every conversation check got a typed answer", "the planted cooling-off denial is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional, recordings only) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |

Typed, WebVTT or SRT transcripts need no speech model. `diarize` only turns a recording into a transcript.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`; on 48 GB also `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories (`sudo nvidia-ctk runtime configure --runtime=docker`, then restart Docker).
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull fails, build from source once the `decosa-api` source is published: clone it and run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in it (and `docker compose build llm` from its compose file). If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-fi/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Auto Group compliance>"
```

Create `~/decosa-fi/docker-compose.yml` with exactly these services:

```yaml
name: decosa-fi
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; a deal with 4 add-ons needs about 3,000 generated tokens
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_FI_AUDIO: "0"                      # "1" only with the diarize service
      DECOSA_DIARIZE_URL: ""
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` as written (a host bind mount owned by root makes the API fail on `/data/keys.sqlite`). Run `docker compose up -d` and poll `docker compose ps` until both are healthy (the LLM takes 5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here and that is fine).

## 4. Optional: recordings (diarize)

Only if I want transcripts made from F&I recordings: build `services/diarize` from the decosa-api source on a CUDA PyTorch base image (env `DIARIZE_HOST=0.0.0.0 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0`, the `hf-cache` volume, the same GPU, a health check on `GET /health`). Lower `LLM_GPU_UTIL` to 0.80, and set `DECOSA_DIARIZE_URL: http://diarize:8092` and `DECOSA_FI_AUDIO: "1"`. `POST /fi/transcribe` then takes a 16 kHz mono 16-bit WAV (`ffmpeg -i in.m4a -ac 1 -ar 16000 -sample_fmt s16 out.wav`) and returns timed segments with a speech receipt. Pass them to `/fi/check` as `segments`, with `speakers` to name S01 and S02. The fit beside the LLM is an estimate, not measured.

## 5. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"fi-disclosure-record"}' | jq -r .token)
curl -s $API/fi/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"tb4-no-consent-no-cooling-off"}' > /tmp/fi.json
jq '{counts, retain_until, receipts: (.receipts|length), statuses: [.receipts[].status] | unique}' /tmp/fi.json
jq -r '.checklist[] | "\(.status)\t\(.answer)\t\(.at // "")\t\(.id)"' /tmp/fi.json
jq -r '.mismatches[] | "\(.status)\t\(.kind)\t\(.what)"' /tmp/fi.json
jq '{record}' /tmp/fi.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

This sample is a $12,900 used car. The F&I manager says there is no cooling-off period, GAP is left undecided, and a paint and fabric package on the jacket is never discussed. Pass if:
- `cancel_right_misstated` is `flag` with a time;
- the mismatches include `never_discussed` for "Paint and fabric protection" and `no_clear_yes` for GAP;
- every receipt has `"status": "attested"`;
- the record verifies (`ok: true`), and `retain_until` is two years from today.

Then try a pasted transcript with your own jacket and consent:

```bash
curl -s $API/fi/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{
  "transcript": "[00:00] FINANCE MANAGER: This office is recorded. Okay with you?\n[00:03] CUSTOMER: Yes.\n[00:05] FINANCE MANAGER: The bank requires the service contract. It is nineteen ninety-five.\n[00:11] CUSTOMER: Okay.",
  "jacket": {"vehicle": {"used": true, "price": 22000}, "addons": [{"name": "Service contract", "kind": "service_contract", "price": 1995}]},
  "consent": {"all_parties": true, "method": "verbal_on_recording"}}' | jq -r '.checklist[] | select(.basis=="conversation") | "\(.status)\t\(.answer)\t\(.id)"'
```

Pass if `addon_required:a1` is `flag`, with the quote at 00:05. Without `consent`, the API answers 400 and cites Penal Code 632.

## 6. Point the app at the local API

- Base URL `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to `DECOSA_CORS_ORIGINS`.
- `POST /fi/check` returns JSON, or streams Server-Sent Events with `Accept: text/event-stream`. The jacket shape and the check ids are listed at `GET /fi/info`.
- The server stores nothing. Keep each signed record (JSON) with the deal jacket for two years. Anyone can re-check it with `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts to the hosted Decosa
API, and those prompts contain the transcript and the jacket's add-ons. Never use it for real deals; at most, use it for synthetic training material.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

3 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
What the F&I manager said about each add-on, quoted with its time, checked against the deal jacket under California's CARS Act.
Who it's for
Teams in sales and marketing and finance and insurance.
Where it runs
Self-host for real deals (hosted demo: synthetic role-plays only)
Key numbers
  • 31/32 Per-check accuracy, test B (held out), run 1 (held out, n = 32)
  • 8/8 Planted answers found, test B (held out), run 1 (held out, n = 8)
  • 1 of 14 False alarms on clean calls, test B (held out), run 1 (held out, n = 14)
  • 3.3 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
MOSS-Transcribe-Diarize · Qwen3.8-27B
Where
Self-host for real deals (hosted demo: synthetic role-plays only)
Checks
Receipt per model call; every yes backed by a quote found in the transcript; signed hash-chained record with a retention date
Output
Signed record or verdict · Structured data
Data
Personal data · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Can a recording satisfy the CARS Act disclosures?

No. The Act's add-on and payment disclosures must be in writing. The tool reads them from the deal jacket and checks whether what was said agrees or undercuts them.

What does it check in the conversation?

For each add-on, whether it was said to be optional and whether it was presented as required, each yes, no or unclear with a quote and its time. It also checks recording consent, the total of payments, the lower-payment warning and, on used cars at $50,000 or less, the 3-day right to cancel.

Is recording the F&I office legal in California?

Recording a confidential conversation needs the consent of every party (Penal Code 632(a)). Every run needs a statement that all parties agreed; it goes into the signed record, and a check looks for the consent on the recording itself.

How was it tested?

On 17 synthetic F&I role-plays written by the building agent. On the held-out set it found 8 of 8 planted answers with 1 of 14 false alarms on clean calls (0 on a repeat). It has not been measured on real F&I recordings.

Does it store our deals?

No. Transcripts, jackets and records live in memory for the request; you keep the signed record for two years (1784.44(a)). Self-hosted, nothing leaves the box.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Auto F&I disclosure record

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.