Review a collections call
A yes, no or unclear answer to each Regulation F and X call question, with the quote and its time, plus call-log breaches computed from your log.
Built on: Speaker diarization, Typed judgment, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-log maths and the record need no GPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the collections and servicing call qa API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- collections-call-qa
Use the hosted API
# Decosa collections and servicing call QA: use the hosted API
You are wiring Decosa's call-QA check into this project. It takes the transcript of a recorded debt-collection or
mortgage-servicing call and the account's call log.
- **On the call**, it gives typed answers (`yes`, `no`, `unclear` or `na`), each with a quote and its time. It checks the
debt-collector disclosure, identity before the debt, third-party disclosure, false or threatening statements, a stop
request or dispute brushed aside, and loss-mitigation options on a delinquent mortgage.
- **From the log**, it computes in code the Regulation F call-frequency and calling-hours presumptions, calls after a stop
request or a lawyer, and Regulation X live contact by day 36.
It returns a signed, hash-chained record to keep with the recording for three years. Every model call has a signed
receipt. Use only what is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic material only.** Real calls carry Social Security numbers and account numbers and belong
on a self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- **Consent.** Some states require every party's consent before a call is recorded. Collect a consent statement in the
UI and send it with every check; the API refuses a check without it.
- This is QA triage, not a compliance verdict. Never label a call or an account "compliant".
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "collections-call-qa"}` returns
`{"token", "expires_at", "budget"}`.
- The demo allows a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz) and a token allowance per session (the `budget` in the session response), which is about nine checks.
- Over a limit you get HTTP 429 with `Retry-After`.
- A demo token runs one check at a time (409 otherwise).
3. A check needs about 700 generated tokens plus 170 per typed question, about 2,100 for a call (402 otherwise).
## Endpoints
- `POST /collections/check` (token). Body: `{"transcript": "[00:05] AGENT: ...", "call_log": {...}, "consent": {"recorded_lawfully": true, "method": "all_parties_notified_on_call"}, "speakers"?: {"S01": "Agent"}, "title"?: "..."}`.
- Give exactly one of:
- `transcript`: text, one utterance per line starting with a time like `[01:23]` and a speaker; WebVTT and SRT also work;
- `segments`: `[{start, end, speaker, text}]` in seconds, as a diarizer returns them;
- `sample_id`: from `/collections/samples`, which brings its own call log and consent.
- `call_log`: `{account: {id?, kind: "consumer_debt"|"mortgage", debt_collector: true|false, mortgage?: {missed_due_dates: ["YYYY-MM-DD"], principal_residence?, small_servicer?, bankruptcy?, fdcpa_cease?}}, consumer: {time_zones: ["America/Chicago"], numbers?: {"+1-312-555-0100": {tz?, kind?}}}, calls: [{id, at: "2026-03-02T08:10:00-06:00", number?, direction?: "outbound"|"inbound", party?: "consumer"|"third_party"|"attorney"|"unknown", outcome: "conversation"|"voicemail"|"no_answer"|"not_connected", prior_consent?}], events?: [{kind: "stop_calling"|"attorney"|"prior_consent"|"cease_written"|"dispute_written", at, number?}], recorded_call: "<call id>"}`.
- Use one log per debt.
- Every time needs a UTC offset.
- List every time zone the file points to.
- `debt_collector` is false for a creditor or servicer collecting its own loan; then the Regulation F items come back `na`.
- `consent.method`: `all_parties_notified_on_call`, `written`, `one_party_state` or `other`. `recorded_lawfully` must be `true`.
- Limits: 400 calls, 60 events, 800 transcript lines, 80,000 characters, 512 KB of JSON.
- JSON response by default:
- `checklist`: `[{id, rule, title, basis: "call"|"call_log", answer, status: "ok"|"flag"|"review"|"na", reason, line?, at?, t_ms?, quote?, evidence?: "quote"|"context", probability?, receipt_ids?, calls?: [{id, local, outcome, outside?, note?}]}]`
- `said`: what the called person asked for or said, each with its line and time
- also `log_events`, `counts`, `call: {initial_communication, …}`, `retain_until`, `record`, `record_check`, `receipts`, `law`, `note`
- Call check ids (`basis: "call"`): `recording_disclosed`, `mini_miranda`, `right_party_verified`,
`debt_before_verification`, `third_party_disclosure`, `false_misleading`, `stop_request_refused`, `dispute_dismissed`,
and on a delinquent mortgage `loss_mit_informed` and `loss_mit_misstated`.
- Log item ids (`basis: "call_log"`): `freq_7in7`, `freq_after_conversation`, `calling_hours`, `after_stop`,
`after_attorney`, `after_cease`, `after_written_dispute`, `call_pattern`, `third_party_repeat` and `live_contact_36`.
- With `Accept: text/event-stream` (or `"stream": true`), the events are:
- `ready`;
- a `receipt` for each model call, and a `check` for each answer as it completes;
- then `log`, `result`, `budget` and `done`.
- `POST /collections/log` (token, no model call) `{"call_log": {...}}` → `{items, stats, events, counts}`: the log findings alone.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, checks, first_bad}`.
- `GET /collections/info`, `GET /collections/samples` and `GET /attest/signing-key` need no token.
`POST /collections/transcribe` is self-host only (503 here).
## Example: check a call and print what needs a person's eyes (Python, `pip install httpx`)
```python
import httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"transcript": open("call-transcript.txt").read(), "call_log": json.load(open("call-log.json")),
"consent": {"recorded_lawfully": True, "method": "all_parties_notified_on_call"}}
r = httpx.post(f"{API}/collections/check", json=body, headers=H, timeout=300)
r.raise_for_status()
js = r.json()
for x in js["checklist"]:
if x["status"] in ("flag", "review"):
where = f" ({x['evidence']} at {x['at']}: \"{x.get('quote', '')[:100]}\")" if x.get("at") else ""
calls = f" calls {', '.join(c['id'] for c in x.get('calls') or [])}" if x.get("calls") else ""
print(f"{x['status'].upper()}: {x['title']} {x['answer']}{where}{calls}")
json.dump(js["record"], open(f"call-qa-record-{js['retain_until']}.json", "w")) # keep with the recording until retain_until
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa collections and servicing call QA: run it yourself (containers)
You are setting up the Decosa collections and servicing call QA on this machine, so call recordings and call logs never
leave it. It checks each recorded collection or servicing call against Regulation F and Regulation X. On the call it gives
typed answers with quotes and times. From the call log it computes the frequency and calling-hours findings in code. It
seals a signed record to keep with the recording for three years. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/collections-call-qa.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py collections-call-qa` (the api image carries the same bundle under /app/rehearsal/collections-call-qa/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py collections-call-qa --bundle collections-call-qa.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least six call checks ran", "every call check got a typed answer", "the planted sheriff threat is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
- set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
- keep its data on a named volume;
- bind every port to 127.0.0.1.
For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
`DECOSA_COLLECTIONS_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/collections/info` lists the rules with their eCFR sections.
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"collections-call-qa"}`, then
`POST /collections/check {"sample_id": "tb1-showcase"}`. Expect:
- `false_misleading` as `flag`, with a time;
- `freq_7in7` as `flag`, with call `c8`;
- `calling_hours` as `flag`, with call `c1`;
- every receipt `attested`.
Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.
Some states require every party's consent before a call is recorded. This is QA triage for a person to review, not a
compliance verdict and not legal advice.
Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds call
recordings or account data. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (61.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (61.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"collections-call-qa"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py collections-call-qa
Download the mock-data bundle (3 KB, 12 checks)expected.json
Nine script lines (twelve diarized segments) of a role-played first collection call and the account call log. The agent threatens the sheriff; the log has 8 counted calls in 7 days and one call at 7:10 a.m. in the mobile number's time zone (Denver). All three must be flagged, the call-log findings must also come back with no model call, and the signed record must verify with a three-year retention date.
What the rehearsal checks
- at least six call checks ran
- every call check got a typed answer
- the planted sheriff threat is flagged
- the threat flag points at a time in the recording
- 8 counted calls in 7 days are flagged
- the 7:10 a.m. call in the Denver zone is flagged
- the log-only route flags the same two findings with no model call
- a check without a consent statement is refused
- the record keeps the recording three years after the call
- the signed record verifies
- a record with its flag count changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic role-play written from 12 CFR part 1006: Harbor Ridge Recovery, Brightwater Card Company and every person are fictional; phone numbers are in the 555-01xx fiction range. Audio spoken by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-collections-demo; the segments are the diarizer output with its errors kept. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa collections and servicing call QA: run it yourself (containers)
You are setting up the Decosa collections and servicing call QA on this machine, so call recordings and call logs never
leave it. It checks each recorded collection or servicing call against Regulation F and Regulation X. On the call it gives
typed answers with quotes and times. From the call log it computes the frequency and calling-hours findings in code. It
seals a signed record to keep with the recording for three years. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/collections-call-qa.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py collections-call-qa` (the api image carries the same bundle under /app/rehearsal/collections-call-qa/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py collections-call-qa --bundle collections-call-qa.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least six call checks ran", "every call check got a typed answer", "the planted sheriff threat is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit, then check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. On the `api` service:
- set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1` and `DECOSA_LLM_MODEL=qwen3.8-27b`;
- keep its data on a named volume;
- bind every port to 127.0.0.1.
For recordings, also keep the `diarize` service and set `DECOSA_DIARIZE_URL=http://diarize:8092` and
`DECOSA_COLLECTIONS_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check; the first start
downloads about 20 GB of weights.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/collections/info` lists the rules with their eCFR sections.
`GET /attest/signing-key` shows this box's public key. Show me the key: it is what others pin to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"collections-call-qa"}`, then
`POST /collections/check {"sample_id": "tb1-showcase"}`. Expect:
- `false_misleading` as `flag`, with a time;
- `freq_7in7` as `flag`, with call `c8`;
- `calling_hours` as `flag`, with call `c1`;
- every receipt `attested`.
Then `POST /record/verify {"record": <record>}` should give `ok: true`.
6. Report back: the public key and key id, the smoke-test results, and how long the check took.
Some states require every party's consent before a call is recorded. This is QA triage for a person to review, not a
compliance verdict and not legal advice.
Off by default. Joining as a provider serves other people's requests on this GPU. Never do it on a box that holds call
recordings or account data. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsCollections and servicing call QA on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Standard · the hosted demo, one 96 GB card: what changesuses estimates
- Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Extraction: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
- Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Collections and servicing call QA, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Collections and servicing call QA on my hardware Fetch https://decosa.ai/prompts/collections-call-qa-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=collections-call-qa) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Extraction: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. - Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/collections-call-qa-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 4.2 s · ~$0.003 per run · 9 receipts
Loading the nightly status…
Self-host: verified 26 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after
Measured cost to run: about $0.35 per 100 calls (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The assembly prompt's smoke tests ran against the already-running local Qwen3.8-27B vLLM (127.0.0.1:8114, network_mode host instead of the compose llm service): tb1-showcase flagged the sheriff threat at 00:27, 7-in-7 (c8) and calling hours (c1), 9 attested receipts, record verified, 3.4 s; the pasted 21:30 New York call flagged the threat and the hours; no consent gave 400. One prompt bug found and fixed: the log-only step fed back the normalised call log, which the API does not accept as input. Model-server startup and the diarizer path were not re-run.
Known limits (5)
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this vertical when the branch merges.
- Measured on 17 synthetic role-plays written by the building agent, with clean TTS audio. Not measured on real collection calls (accents, Spanish, long calls, voicemails) or with an independent reviewer's labels.
- One systematic miss: an agent who names themselves a debt collector before confirming who answered is not flagged as revealing the debt before identity (t6 on every run).
- The call-log findings are the rule's presumptions computed from the log given. Letters, texts, emails, consent given elsewhere and an attorney's response are outside it unless added as events. Days are counted in the consumer's first time zone.
- Audio intake (/collections/transcribe) is self-host only. The hosted demo's audio samples were transcribed on our server and are bundled.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-log maths and the record need no GPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the collections and servicing call qa API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
A recorded collection or servicing call and its call log, checked against Regulation F and X: each answer quoted with its time, the call counts computed in code.
For debt-collection agencies, debt buyers and mortgage servicers that record calls. Give it the call (audio on your own box, or a transcript) and the account's call log. On the call it answers yes, no or unclear, each with a quote and its time: the debt-collector disclosure, identity confirmed before the debt came up, the debt disclosed to a third party, false or threatening statements, a stop request or dispute brushed aside, and for a delinquent mortgage whether loss-mitigation options were mentioned or misstated. From the log, in code: more than 7 calls in 7 days, a call within 7 days after a conversation, calls before 8 a.m. or after 9 p.m. in every time zone the file points to, calls after a stop request or a lawyer, and live contact by day 36. Everything is sealed in a signed, hash-chained record to keep with the recording for three years. It is QA triage for a compliance reviewer, not a compliance verdict.
- Deployment
- Self-host first
- Regulatory
- Checked 26 Sep 2026 against eCFR (text current to 24 Sep 2026). Regulation F, 12 CFR part 1006 (in force since 30 Nov 2021), binds debt collectors as defined in 1006.2(i): third-party collectors and debt buyers, and a mortgage servicer only for loans it obtained in default. Encoded: the debt-collector disclosure in the initial and later communications (1006.18(e)); no communication about the debt with third parties (1006.6(d)(1); location calls must not say a debt is owed, 1006.10(b)); no false, deceptive or misleading representations or threats (1006.18); no calls through a medium the person asked not to use, e.g. "stop calling" (1006.14(h), comment 14(h)(1)-3.i); the presumption of a violation for more than seven calls in seven consecutive days, or a call within seven days after a telephone conversation, per person and per debt, with the (b)(3) exclusions (1006.14(b)(2)); the 8 a.m.-9 p.m. local-time presumption, convenient in every location the collector's information points to (1006.6(b)(1)(i), comment 6(b)(1)(i)-2); no calls to a consumer known to be represented by an attorney (1006.6(b)(2)); and keeping each call recording for three years after the call (1006.100(b)). Oral disputes are recorded but the stop-collection duty in 1006.38(d)(2) applies only to written disputes in the validation period. Regulation X, 12 CFR 1024.39-1024.41 (as amended at 90 FR 20792, 16 May 2025, which removed the expired COVID-19 live-contact provisions): live contact with a delinquent borrower by day 36, and informing the borrower about loss-mitigation options if appropriate (1024.39(a)); accurate information from assigned personnel (1024.40(b)). They apply to a principal residence and not to small servicers (1024.30). Right-party verification is a practice check that supports 1006.6(d), not a rule of its own. Recording consent is state law (e.g. California Penal Code 632). QA triage, not legal advice.
Text description
A recorded collection or servicing call (self-host only) goes through MOSS-Transcribe-Diarize to a timed, speaker-labelled transcript; a typed, WebVTT or SRT transcript can be given instead. With it come the account's call log (every call with its time, number and result, the consumer's time zones, which call was recorded) and a recording-consent statement. Qwen3.8-27B extracts what the called person asked for (stop calling, a lawyer, call back, a dispute) and answers one typed question per QA check, each with a quote that must be found in the transcript. Requests heard on the call are placed on the log at their time. Code computes the Regulation F call-frequency presumptions, calling hours in every time zone on file, calls after a stop request, and Regulation X live contact by day 36. Outputs: the QA checklist, the call-log findings and a signed hash-chained record with its retention date (the call date plus three years). Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.
At a glance
- What it checks
- On each recorded call: the debt-collector disclosure, identity before the debt, third-party disclosure, false or threatening statements, a stop request or dispute brushed aside, and loss-mitigation options on a delinquent mortgage. From the call log: the 7-in-7 and 7-days-after-a-conversation presumptions, calling hours in every time zone on file, calls after a stop request or a lawyer, and live contact by day 36.
- What it does not do
- It gives no compliance verdict. It does not check validation notices, letters, emails or texts, credit reporting, time-barred-debt disclosures, Regulation X written notices or loss-mitigation application handling beyond what is said on the call. The log findings are presumptions from the log you give.
- Consent
- Every run needs a statement that the call was recorded with the consent the law requires (some states need every party's). The statement goes into the signed record, and a check looks for the recording notice on the call.
- Data retention
- Nothing on the server. Transcripts, call logs and records live in memory for the request, and logs carry counts and timings only. You keep the signed record with the recording for three years (12 CFR 1006.100(b)).
- What leaves the box (hosted demo)
- The transcript goes to Qwen3.8-27B through our gateway, which is Decosa-operated. The call log is processed in code and not sent to the model. The gateway's receipts hold hashes, not text. Self-hosted, nothing leaves.
- Model calls per call
- One extraction plus 8 typed checks (10 for a mortgage). The call-log findings need no model: POST /collections/log returns them on their own.
- Typical run cost
- A fraction of a cent at the gateway list price for a short call (measured on the smoke sample). Each run shows its own measured cost.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 48 GB card
The same checks on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.
- Models
- MOSS-Transcribe-Diarize 0.9B
- Gemma 4 26B A4B (instruction-tuned)
- Hardware
- 1x L40S or RTX 6000 Ada 48 GB (not measured)
- Quality evidence
- typed checks correct / planted problems foundnot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
- In the hosted demo
Standard
the hosted demo, one 96 GB card
Qwen3.8-27B extracts and answers every typed check; MOSS-Transcribe-Diarize turns recordings into transcripts; the call-log maths is code. Every model call receipted.
- Models
- Qwen3.8-27B (NVIDIA NVFP4)
- MOSS-Transcribe-Diarize 0.9B
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB
- Quality evidence
- held-out test B, 4 synthetic calls: typed checks correct / planted problems found / call-log findings32/33 / 3/3 / 11/11decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route, run 1; data, labels and prompts written by the building agent, prompts frozen on a separate 3-call dev split
- test set, 10 calls: typed checks correct / planted found / call-log findings81/83 / 11/12 / 33/33 (repeat run the same, plus one unlabelled log flag)decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route
- false alarms on clean calls (flag or review)test 0/48; test B 2/21 (repeat 1/21)decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route
- 6 calls on real speech-recognition transcripts of synthetic audio: checks / planted / log findings / cited time within 3 s48/49 / 8/9 / 17/17 / 20/21decosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route (asr runs)
- real collection calls reviewed by a compliance QA leadnot measured yet
- Latency
- measured on our server under a shared gateway: a few seconds per call
- Verification
- Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
Best
DeepSeek-V4-Flash on two more cards
A larger model for long calls and hard cases.
- Models
- MOSS-Transcribe-Diarize 0.9B
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- Hardware
- 2x RTX PRO 6000 96 GB
- Quality evidence
- typed checks correct / planted problems foundnot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
- Needs more compute
Wanted: the best setup
two large judges from different families
DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Call recordings stay on your own hardware, never on community providers. Not served yet.
- Models
- MOSS-Transcribe-Diarize 0.9B
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- GLM-5.3-Flash
- Hardware
- Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
- Quality evidence
- typed checks correct / planted problems foundnot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| After the session | ||||
Extraction (what the called person asked for or said, with the line) and one typed yes/no/unclear check per QA questionQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | Standard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.9BProof: partialSelf-host only | LiteStandardBestWanted | 0.9B | Proof: partialSelf-host only | |
| ||||
Lite tier: the same extraction and checks on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 25.2B (3.8B active)No proof yetSelf-host only | Lite | 25.2B (3.8B active) | No proof yetSelf-host only | |
| ||||
Best tier: a larger model for long calls and hard casesDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | BestWanted | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
| Other | ||||
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) 321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only | Wanted | 321B (18B active) · about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- decosa F&I disclosure record (vertical 46) (opens in a new tab)AGPL-3.0-or-later (decosa-api)
The call-record engine: transcript intake, quote matching and answer parsing; imported, not copied.
- decosa typed-judgment (vertical 24) (opens in a new tab)AGPL-3.0-or-later (decosa-api)
The stated-confidence probability for each typed check.
- decosa record (vertical 07) and POST /record/verify (opens in a new tab)AGPL-3.0-or-later (decosa-api)
The hash chain, signed checkpoints and the Ed25519-signed record; anyone can re-check it.
- Regulation F (12 CFR part 1006) and its official interpretations, eCFR (opens in a new tab)public law
The source of every Regulation F rule; each checklist item links its section.
Early intervention, continuity of contact and loss mitigation for mortgage servicers.
Local times for the calling-hours check, including daylight-saving changes.
- cryptography (Python) (opens in a new tab)Apache-2.0 OR BSD-3-Clause
Ed25519 signatures on the record and the speech receipts.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:0.1.0Transcript and call-log intake, the checks, the call-log maths, signing and the HTTP API (/collections/*). No GPU. Binds 127.0.0.1 by default.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.
- decosa-diarize:8092
Optional, self-host only: MOSS-Transcribe-Diarize for transcripts made from recordings (DECOSA_COLLECTIONS_AUDIO=1). No published image yet; built from services/diarize.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server; the diarizer runs on the other card there.
- 1x L40S / RTX 6000 Ada 48 GB
Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).
- CPU only Fits
The call-log findings (POST /collections/log), transcript parsing, the signed record and verification need no GPU; the call checks need the model.
Latency per lane
- check of a 7-10 line call with its call log, busy shared gateway4.2 s
Measureddecosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route (medians 4.2-5.2 s per call across the test and test B runs; slowest 16.3 s)
- check of a diarized audio transcript, busy shared gateway4.0 s
Measureddecosa-api docs/evals/collections-call-qa.md, measured on our server 2026-09-26, gateway route (asr runs, medians 4.0 and 4.5 s)
- call-log findings only (POST /collections/log)20 ms
Estimateestimate: code only, no model call
- transcribe a 32-51 s recording (self-host)3.0 s
Measuredmeasured on our server 2026-09-26, MOSS-Transcribe-Diarize on an RTX PRO 6000
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa collections and servicing call QA on this machine
You are setting up a self-hosted call-QA checker on this Linux machine for a debt-collection agency or a mortgage servicer.
It takes the transcript of a recorded collection or servicing call and the account's call log.
- **On the call**, each answer is yes, no or unclear, with a quote and a time. It checks:
- the debt-collector disclosure;
- whether identity was confirmed before the debt came up;
- whether the debt was disclosed to a third party;
- false or threatening statements;
- whether a stop request or a dispute was brushed aside;
- for a delinquent mortgage, whether loss-mitigation options were mentioned.
- **From the call log**, computed in code:
- the Regulation F call-frequency presumptions (12 CFR 1006.14(b)(2));
- calling hours in every time zone on file (1006.6(b)(1)(i));
- calls after a stop request or a lawyer;
- Regulation X live contact by day 36 (1024.39(a)).
- It seals everything in a signed, hash-chained record to keep with the recording for three years (1006.100(b)).
- With the optional diarizer, it turns a recording into a timed, speaker-labelled transcript.
Work step by step. Show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Collection calls carry Social Security numbers, dates of birth and account numbers. Keep everything on this machine: the
model route stays local (`direct`), and nothing goes to a hosted service.
- Some states require every party's consent before a call is recorded (for example California Penal Code 632). Each
check needs a consent statement, and the statement goes into the record.
- This is QA triage for a person to review. It is not a compliance verdict and not legal advice.
Repeat these points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/collections-call-qa.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py collections-call-qa` (the api image carries the same bundle under /app/rehearsal/collections-call-qa/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py collections-call-qa --bundle collections-call-qa.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "at least six call checks ran", "every call check got a typed answer", "the planted sheriff threat is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional, recordings only) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |
Typed, WebVTT or SRT transcripts need no speech model. The call-log findings need no model at all (`POST /collections/log`).
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`; on 48 GB also
set `LLM_MAX_LEN=32768`. Not measured.
- Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories (`sudo nvidia-ctk runtime configure --runtime=docker`, then restart Docker).
3. Confirm about 60 GB of free disk.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file.
1. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
2. If a pull fails, build from source once the `decosa-api` source is published. Clone it, then run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in it, and
`docker compose build llm` from its compose file.
3. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-cq/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Recovery compliance>"
```
Create `~/decosa-cq/docker-compose.yml` with exactly these services:
```yaml
name: decosa-cq
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "200000" # per session; a call needs about 2,100 generated tokens
DECOSA_SESSION_TTL_S: "28800"
DECOSA_COLLECTIONS_AUDIO: "0" # "1" only with the diarize service
DECOSA_DIARIZE_URL: ""
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volume `decosa-data` as written: a host bind mount owned by root makes the API fail on
`/data/keys.sqlite`.
1. Run `docker compose up -d`.
2. Poll `docker compose ps` until both services are healthy. The LLM takes 5-10 minutes the first time.
3. `curl -s localhost:8445/healthz` should show `"llm": true`. `asr` is false here, and that is fine.
## 4. Optional: recordings (diarize)
Do this only if I want transcripts made from call recordings.
1. Build `services/diarize` from the decosa-api source on a CUDA PyTorch base image. Give it the env
`DIARIZE_HOST=0.0.0.0 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0`, the `hf-cache` volume, the same GPU, and a health
check on `GET /health`.
2. Lower `LLM_GPU_UTIL` to 0.80. On the api, set `DECOSA_DIARIZE_URL: http://diarize:8092` and
`DECOSA_COLLECTIONS_AUDIO: "1"`.
3. `POST /collections/transcribe` then takes a 16 kHz mono 16-bit WAV
(`ffmpeg -i in.m4a -ac 1 -ar 16000 -sample_fmt s16 out.wav`) and returns timed segments with a speech receipt.
4. Pass the segments to `/collections/check` as `segments`, with `speakers` to name S01 and S02.
The fit beside the LLM is an estimate, not measured.
## 5. Smoke test
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"collections-call-qa"}' | jq -r .token)
curl -s $API/collections/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"sample_id":"tb1-showcase"}' > /tmp/cq.json
jq '{counts, retain_until, receipts: (.receipts|length), statuses: [.receipts[].status] | unique}' /tmp/cq.json
jq -r '.checklist[] | "\(.status)\t\(.answer)\t\(.at // "")\t\(.id)\t\([.calls[]?.id] | join(","))"' /tmp/cq.json
jq '{record}' /tmp/cq.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```
In this sample, a fictional agency makes its first contact after 8 calls in 7 days. One call was at 8:10 a.m. Central to
a mobile with a Denver number. On the call the agent threatens the sheriff. Pass if:
- `false_misleading` is `flag` with a time;
- `freq_7in7` is `flag` with call `c8`;
- `calling_hours` is `flag` with call `c1`;
- every receipt has `"status": "attested"`;
- the record verifies (`ok: true`), and `retain_until` is `2029-03-08`.
Then try the call log on its own (no model: `freq_7in7` and `calling_hours` flag again), and a pasted transcript with your own log and consent:
```bash
curl -s $API/collections/log -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"tb1-showcase"}' | jq -r '.items[] | "\(.status)\t\(.id)"'
curl -s $API/collections/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{
"transcript": "[00:00] AGENT: This call is recorded. Is this Alex Doe?\n[00:04] CONSUMER: Yes.\n[00:06] AGENT: We will have you arrested if you do not pay today.",
"call_log": {"account": {"kind": "consumer_debt", "debt_collector": true}, "consumer": {"time_zones": ["America/New_York"]},
"calls": [{"id": "c1", "at": "2026-03-02T21:30:00-05:00", "outcome": "conversation"}], "recorded_call": "c1"},
"consent": {"recorded_lawfully": true, "method": "all_parties_notified_on_call"}}' | jq -r '.checklist[] | "\(.status)\t\(.answer)\t\(.id)"'
```
Pass if `false_misleading` and `calling_hours` (21:30 in New York) are `flag`. Without `consent`, the API answers 400.
## 6. Point the app at the local API
- The base URL is `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to
`DECOSA_CORS_ORIGINS`.
- `POST /collections/check` returns JSON, or streams Server-Sent Events with `Accept: text/event-stream`.
`GET /collections/info` lists the call-log fields and the check ids.
- Map your dialer export onto the call log:
- one log per debt;
- times with a UTC offset;
- `outcome` is `conversation`, `voicemail`, `no_answer` or `not_connected` (busy, not in service);
- every time zone the file points to goes in `consumer.time_zones`, or as `tz` on a number;
- add letters and written notices as `events`.
- The server stores nothing. Keep each signed record (JSON) with the recording for three years. Anyone can re-check it
with `POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set
`DECOSA_TRUSTED_PROXIES`.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain the call transcript, to the hosted Decosa
API. Never use it for
real calls; at most, use it for synthetic training material.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
11 laws, rules and guidance pages cited; 10 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A recorded collection or servicing call and its call log, checked against Regulation F and X: each answer quoted with its time, the call counts computed in code.
- Who it's for
- Teams in finance and insurance and compliance and trust.
- Where it runs
- Self-host for real calls (hosted demo: synthetic role-plays only)
- Key numbers
- 32/33 Per-rule accuracy, test B run 1 (written after test run 1) (held out, n = 33)
- 81/83 Per-rule accuracy, test run 1 (test split, n = 83)
- 11/12 Planted answers found, test run 1 (test split, n = 12)
- 4.2 s Median end-to-end run, hosted (QA sweep 2026-09-26)
- Models
- MOSS-Transcribe-Diarize · Qwen3.8-27B
- Where
- Self-host for real calls (hosted demo: synthetic role-plays only)
- Checks
- Receipt per model call; every yes backed by a quote found in the transcript; call-log findings computed in code; signed hash-chained record with a retention date
- Industry
- Finance and insurance · Compliance and trust
- Runs
- Self-host
- Output
- Signed record or verdict · Structured data
- Data
- Personal data · Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Speaker diarization · Typed judgment · Signed record
Questions people ask
What does it check on a recorded call?
Yes, no or unclear, each with a quote and its time: the debt-collector disclosure, identity confirmed before the debt came up, the debt disclosed to a third party, false or threatening statements, a stop request or dispute brushed aside, and for a delinquent mortgage whether loss-mitigation options were mentioned or misstated.
How does it count the 7-in-7 rule?
In code from the call log, per person and per debt, with the 1006.14(b)(3) exclusions, plus calls within seven days after a telephone conversation. The findings are the rule's presumptions from the log you give; letters, texts and consent given elsewhere are outside it unless added as events. It matched all 61 hand-labelled log findings.
How are calling hours checked when a consumer has two time zones?
A call must fall between 8 a.m. and 9 p.m. in every time zone the file points to, following comment 6(b)(1)(i)-2.
Is it a compliance verdict?
No. It is QA triage for a compliance reviewer. It does not check validation notices, letters, emails or texts, credit reporting or time-barred-debt disclosures.
How accurate is it?
On 17 synthetic scripted calls: 81 of 83 per-rule answers right on the test set, 32 of 33 on a set written afterwards, and 48 of 49 on ASR transcripts of TTS audio. Real calls (accents, Spanish, long calls, voicemails) are not measured, and one systematic miss is documented.
How does it fit the three-year recording rule?
Nothing is kept on the server. You keep the signed, hash-chained record with the recording for three years (12 CFR 1006.100(b)). Every run also records a statement that the call was recorded with the consent the law requires.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Collections and servicing call QA
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…