Quote a job from a walkthrough video
A draft scope and quote where each line cites a moment in the video, is marked seen, said or assumed, and is priced only from your own price sheet.
Built on: Video understanding, Speaker diarization, Numeric grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the walkthrough-to-quote API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on Qwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, pricing and the record on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- walkthrough-to-quote
Use the hosted API
# Decosa Walkthrough-to-quote: use the hosted API
You are wiring Decosa's Walkthrough-to-quote into this project (a field-service app, a quoting tool or an estimator's
back office). It takes a narrated phone video of a job walk and the business's own price sheet and returns a draft
line-item scope and quote: each line cites a time range and a keyframe and is marked seen, said, seen and said, or
assumed; quantities say where they came from (said, estimated from the video as a range, measured, or missing); prices
come only from the sheet, in integer cents, with a signed record of the arithmetic. Use only what is listed below. If you
need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- It is a draft for the estimator. Never send it to a customer as a quote without a person measuring and signing; never
add prices that are not on the business's sheet.
- Send videos you have the right to upload, filmed with the customer's knowledge. Videos of homes are personal data;
confidential ones belong on a self-hosted box.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, kept in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "walkthrough-to-quote"}` returns
`{"token", "expires_at", "budget"}`. Sessions per IP are limited (HTTP 429 with `Retry-After`); a demo token runs one
draft at a time (409).
3. A draft needs about 9,900 generated tokens of budget before it starts (402 otherwise); a one-minute walk uses far less.
## Endpoints
- `POST /walkthrough/draft` (token). Body: `{"video_b64": "<MP4, MOV or WebM as base64>" | "sample": "<id>", "price_sheet_csv"?: "code,description,unit,unit_price\n...", "title"?: "...", "trade"?: "...", "transcript"?: [{"start": 2.1, "end": 3.8, "text": "..."}], "stream"?: true}`.
- Videos up to 64 MB and 10 minutes (watched in 2-minute parts past 150 s). Units on the sheet: sq ft, sq yd, lin ft,
each, room, day, hour. Without `transcript`, the narration is transcribed on Decosa's hosted service (a model-call receipt).
- JSON response: `{run_id, counts, totals, record, report, budget}`. `report.lines`:
`[{n, area, work, basis, range, seen, said, quantity: {unit, low, high, source, method}, pricing: {status, code, unit_price, qty_low, qty_high, amount_low, amount_high, arithmetic}, keyframe: {t_s, sha256, jpeg_b64}}]`;
also `report.conditions`, `report.materials`, `report.measure_first`, `report.quote.totals`, `report.timing_warning`.
- With `Accept: text/event-stream` (or `"stream": true`): `video`, `transcript`, `receipt`, `survey`, `narration`,
`scope`, `said_only`, `quote`, `keyframes`, `timing_warning`, `report`, `done`, `budget`.
- `POST /walkthrough/runs/{run_id}/price` (same token) `{"measurements"?: {"<line n>": 250 | [240, 260]}, "codes"?: {"<line n>": "CODE" | null}, "remove"?: [n]}`
→ re-priced lines and totals and a new signed record; no model call. Runs are kept for one hour.
- `POST /walkthrough/runs/{run_id}/signoff` `{"name", "role", "note"?, "confirm": true}` → `{record, check, totals}`.
- `GET /walkthrough/runs/{run_id}/export?format=md|json|csv|record`, `POST /walkthrough/check-arithmetic`
`{"quote", "price_sheet_csv"}` and `POST /record/verify` `{"record"}` (no token), `GET /walkthrough/info`,
`GET /walkthrough/samples`, `GET /walkthrough/samples/{id}/video|prices`, `GET /attest/signing-key`.
## Example: draft, measure what matters, sign (Python, `pip install httpx`)
```python
import base64, httpx, json, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
video = base64.b64encode(open("walkthrough.mp4", "rb").read()).decode()
sheet = open("price-sheet.csv").read()
r = httpx.post(f"{API}/walkthrough/draft", json={"video_b64": video, "price_sheet_csv": sheet, "title": "14 Elm St repaint"},
headers=H, timeout=900).json()
rep = r["report"]
for ln in rep["lines"]:
p = ln["pricing"]
print(ln["n"], ln["range"], ln["basis"], ln["work"], ln["quantity"]["source"], p.get("amount_low"), p.get("amount_high"))
measured = {}
for m in rep["measure_first"]:
v = input(f"line {m['n']} ({m['work']}): measure {m['measure']}, or Enter to skip: ").strip()
if v:
measured[str(m["n"])] = float(v)
if measured:
httpx.post(f"{API}/walkthrough/runs/{r['run_id']}/price", headers=H, json={"measurements": measured}).raise_for_status()
signed = httpx.post(f"{API}/walkthrough/runs/{r['run_id']}/signoff", headers=H, json={"name": "A. Estimator", "role": "estimator", "confirm": True}).json()
open("quote.csv", "w").write(httpx.get(f"{API}/walkthrough/runs/{r['run_id']}/export?format=csv", headers=H).text)
json.dump(signed["record"], open("quote.record.json", "w")) # anyone can re-check it at /record/verify
```
## Receipts
Every model call on the hosted route gets a gateway-signed receipt; the video calls' request hash covers the video's
sha256 and the sampling settings. The speech to text gets a model-call receipt signed by Decosa's instance key. The
quote record lists them and holds hashes and numbers only.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Walkthrough-to-quote: run it yourself (containers)
You are setting up Decosa's Walkthrough-to-quote on this machine, so videos of customers' homes never leave it. It turns
a narrated job walkthrough and the business's own price sheet into a draft scope and quote: lines that cite a time range
and a keyframe, quantities that say where they came from, prices only from the sheet with the arithmetic in code, and a
signed record. Nothing is sent to Decosa's hosted API. It drafts for the estimator; it does not send quotes.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/walkthrough-to-quote.zip (2.2 MB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py walkthrough-to-quote` (the api image carries the same bundle under /app/rehearsal/walkthrough-to-quote/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py walkthrough-to-quote --bundle walkthrough-to-quote.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the ceiling water stain (shown, never said) is a seen-only line", "the guest-bath ceiling (said, never filmed) is a said-only line", "the bedroom's length is the one said on camera (14 ft)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM; it must run with `--limit-mm-per-prompt '{"image":4,"video":1}'`
and `--media-io-kwargs '{"video":{"num_frames":240,"fps":1}}'`, not `--language-model-only`) and the `api` service.
For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`
and bind every port to 127.0.0.1. For narration, keep a `diarize` service if the compose file has one and set
`DECOSA_DIARIZE_URL=http://diarize:8092`; otherwise send the transcript with each video. Never set the gateway route
on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Smoke test: get a token with `POST /demo/session {"vertical":"walkthrough-to-quote"}` and send
`{"sample": "painter-bedroom-hall", "stream": false}` to `POST /walkthrough/draft`. Expect 6-10 lines, each with a
keyframe; the ceiling water stain as a `seen` line; the guest-bath ceiling as a `said` line with quantity `missing`;
a subtotal range; every receipt `"attested"` (signed by this box). Then `POST /walkthrough/check-arithmetic` with the
report and `GET /walkthrough/samples/painter-bedroom-hall/prices`, and `POST /record/verify` with `record`: both `ok`.
5. Report back: `GET /attest/signing-key` (the public key a reviewer pins), the smoke-test results and how long the draft took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
customers' videos. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVIDIA NVFP4), video input needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (61.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (61.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"walkthrough-to-quote"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py walkthrough-to-quote
Download the mock-data bundle (2.2 MB, 11 checks)expected.json
A 53-second walkthrough of a rendered bedroom and hallway, narrated by a stock synthetic voice. The narrator gives the bedroom's size (twelve by fourteen, eight-foot ceiling), asks for the walls, the trim and both closet doors, points at peeling paint, never mentions a water stain the camera shows on the ceiling, and asks for a guest-bath ceiling that is never filmed. The draft must list the stain as a seen-only line and the guest-bath ceiling as a said-only line, use the size said on camera for the bedroom, cite a keyframe for every line, price only from the painter's sheet with arithmetic that re-computes exactly, and seal a record that verifies and fails once changed.
What the rehearsal checks
- the ceiling water stain (shown, never said) is a seen-only line
- the guest-bath ceiling (said, never filmed) is a said-only line
- the bedroom's length is the one said on camera (14 ft)
- every line has a keyframe
- at least five scope lines
- every quantity is said, estimated with a method, per room or missing: none is made up
- the arithmetic re-computes exactly from the price sheet
- a price sheet without the required columns is refused
- the signed record of the arithmetic verifies
- the record fails once it is changed
- every model call has a signed receipt
Licence: Synthetic: rooms rendered with three.js (MIT) from a scripted scene written for Decosa, narrated by the Kokoro-82M stock voice am_michael (Apache-2.0). No real homes, people or voices. Video, transcript and price sheet CC0; part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa Walkthrough-to-quote: run it yourself (containers)
You are setting up Decosa's Walkthrough-to-quote on this machine, so videos of customers' homes never leave it. It turns
a narrated job walkthrough and the business's own price sheet into a draft scope and quote: lines that cite a time range
and a keyframe, quantities that say where they came from, prices only from the sheet with the arithmetic in code, and a
signed record. Nothing is sent to Decosa's hosted API. It drafts for the estimator; it does not send quotes.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/walkthrough-to-quote.zip (2.2 MB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py walkthrough-to-quote` (the api image carries the same bundle under /app/rehearsal/walkthrough-to-quote/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py walkthrough-to-quote --bundle walkthrough-to-quote.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the ceiling water stain (shown, never said) is a seen-only line", "the guest-bath ceiling (said, never filmed) is a said-only line", "the bedroom's length is the one said on camera (14 ft)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM; it must run with `--limit-mm-per-prompt '{"image":4,"video":1}'`
and `--media-io-kwargs '{"video":{"num_frames":240,"fps":1}}'`, not `--language-model-only`) and the `api` service.
For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`
and bind every port to 127.0.0.1. For narration, keep a `diarize` service if the compose file has one and set
`DECOSA_DIARIZE_URL=http://diarize:8092`; otherwise send the transcript with each video. Never set the gateway route
on this box.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Smoke test: get a token with `POST /demo/session {"vertical":"walkthrough-to-quote"}` and send
`{"sample": "painter-bedroom-hall", "stream": false}` to `POST /walkthrough/draft`. Expect 6-10 lines, each with a
keyframe; the ceiling water stain as a `seen` line; the guest-bath ceiling as a `said` line with quantity `missing`;
a subtotal range; every receipt `"attested"` (signed by this box). Then `POST /walkthrough/check-arithmetic` with the
report and `GET /walkthrough/samples/painter-bedroom-hall/prices`, and `POST /record/verify` with `record`: both `ok`.
5. Report back: `GET /attest/signing-key` (the public key a reviewer pins), the smoke-test results and how long the draft took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
customers' videos. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsWalkthrough-to-quote on GeForce RTX 5090: use the Standard · narration, scope and quote (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB: Not measured. The weights are about 20 GB; a maximum video request is 32k prompt tokens of KV cache.
Standard · narration, scope and quote (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Watches the walkthrough: Qwen3.8-27B (NVIDIA NVFP4), video input. ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
- The narration as timed lines, with a model-ca...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.
- Clip preparation and chunking, keyframes, qua...: decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough). CPU. Runs on CPU (vram_gb 0 in stack.json).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Walkthrough-to-quote, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Walkthrough-to-quote on my hardware Fetch https://decosa.ai/prompts/walkthrough-to-quote-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=walkthrough-to-quote) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · narration, scope and quote (hosted demo) (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Watches the walkthrough: Qwen3.8-27B (NVIDIA NVFP4), video input (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. - The narration as timed lines, with a model-ca...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB - Clip preparation and chunking, keyframes, qua...: decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough), CPU GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVIDIA NVFP4), video input ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/walkthrough-to-quote-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 54 s · ~$0.020 per run · 5 receipts
Loading the nightly status…
Self-host: verified 27 Sep 2026 · Fresh clone of the pre-release branch into a clean directory, docker build of the api image (39 s), the api with a named volume on host networking against the local vLLM (video on) and diarizer, direct route; then torn down.
Measured cost to run: about $0.020 per walkthrough (hosted, 27 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 11 of 11 in 28.6 s; the flooring sample with speech to text inside the container drafted 12 lines (8 priced, $4,037.00-$4,453.50) in 38 s with 4 attested receipts and a model-call receipt for the ASR; the quote record verified; no line, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load).
Known limits (4)
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.
- Measured on six synthetic walkthroughs of rendered rooms, made and labelled by the building agent; real phone footage is not measured.
- Video area estimates are often far off on held-out clips (2 of 5 inside the range); the draft marks them estimated and ranks them in measure first.
- Runs vary: the demo painter sample came back with 8 to 10 lines across runs; on one run the ceiling stain was a condition only (since then visible damage always gets a line).
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the walkthrough-to-quote API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on Qwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, pricing and the record on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
A narrated phone walkthrough of a job, turned into cited scope lines and a draft quote priced only from your own price sheet.
The customer or estimator walks the job with a phone and talks. Qwen3.8-27B watches the video and lists the areas and everything visible that matters (damage, hazards, access, fixtures), each with its time; the narration is transcribed and read for requests and spoken measurements; the scope lines cite both and are marked seen, said, seen and said, or assumed. Code works out quantities (said numbers only if the transcript has them, video estimates as ranges with their method, otherwise a request to measure), prices them from your sheet in integer cents and signs the arithmetic. The estimator measures, re-prices and signs; the business sets the final price. For painters, flooring, remodel, restoration, landscaping and moving crews.
- Deployment
- Hosted or self-host
- Regulatory
- Not legal advice; a draft for the estimator, not a contract or a sendable quote. Checked 27 Sep 2026: if the quote leads to a sale agreed at the customer's home, the FTC's Cooling-Off Rule may apply: 16 CFR 429.0(a) defines a door-to-door sale as one where the buyer's agreement is made 'at a place other than the place of business of the seller', at $25 or more at the buyer's residence, and the buyer then has three business days to cancel (https://www.ecfr.gov/current/title-16/chapter-I/subchapter-D/part-429, read via the Cornell LII copy at https://www.law.cornell.edu/cfr/text/16/429.0). Many states also set rules for home-improvement contracts; we did not check them. Recording a conversation: in some US states everyone must consent (California Penal Code 632(a) makes it an offence to record a confidential communication 'without the consent of all parties', https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?sectionNum=632.&lawCode=PEN); film with the customer's knowledge. Video of a home can be personal data; keep people out of frame. The demo walkthroughs are synthetic.
Text description
A walkthrough video and the business's price sheet go to decosa-api. ffmpeg makes a clip with the same timeline and no audio and extracts the audio; MOSS-Transcribe-Diarize turns the narration into timed lines with a model-call receipt. Qwen3.8-27B (Apache-2.0) watches the clip through our gateway: areas with dimension ranges, visible conditions with times, then requests and spoken measurements from the narration, then scope lines citing both. Code computes quantities from said numbers or video ranges, prices them only from the sheet in integer cents and cuts a keyframe per line. The estimator measures, re-prices and signs a record. Self-hosted, everything stays on the box.
At a glance
- What it does
- Turns one narrated walkthrough video and your price sheet into scope lines (each with a time range, a keyframe and a basis: seen, said, seen and said, or assumed), conditions (damage, hazards, access), materials mentioned, quantities with where they came from, prices from your sheet, a subtotal range and the lines to measure first. You enter measurements, re-price, and sign; exports in Markdown, CSV and JSON.
- What it does not do
- It does not price from anything but your sheet (no market prices), add tax, markup or minimum charges, or send a quote to a customer. It does not see behind walls or under floors, check codes or permits, or measure: video quantities are estimates, and on held-out clips areas were often far off. Videos over 4 minutes are watched in 2-minute parts at one frame a second, so brief views can be missed.
- Data retention
- The video and its clip are deleted when the run ends. The report, with its keyframes, is kept in memory for one hour for the token or key that made it. Logs carry run ids and counts, never titles, narration or line text. You keep the exports and the signed record, which holds hashes and numbers only.
- What leaves the box (hosted demo)
- The clip (no audio), the narration text and your price sheet's codes and descriptions go to Qwen3.8-27B through our gateway, which Decosa operates; the audio is transcribed on Decosa's hosted service. The gateway's receipts hold hashes, not content. Self-hosted, nothing leaves.
- Accuracy
- On 4 held-out synthetic walkthroughs: 33 of 33 planted items found, 36 of 37 lines and conditions matched a planted item, the basis right on 26 of 33 (2 of 5 said-only items were wrongly marked seen), counts exact on 11 of 14, but only 2 of 5 video areas and lengths had the truth inside the range (median error 68.7%). Every priced amount re-computed exactly.
- Cost per walkthrough
- A few cents for a short walkthrough at gateway list price (a few model calls, measured on the demo sample), plus speech to text on Decosa's hosted service.
- Output
- A draft quote with lines, time ranges, keyframes, bases, quantities and their method, unit prices and amounts; a decosa.record.v1 record of the arithmetic, re-signed when you re-price and when you sign (verify at /record/verify), and POST /walkthrough/check-arithmetic to re-run the sums.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
scope and quote from the video only
The areas, conditions and scope lines the video shows, each with a time range and a keyframe, priced from your sheet. No speech to text: send a transcript if you have one, otherwise nothing is marked said and spoken measurements are not used.
- Models
- Qwen3.8-27B (NVIDIA NVFP4), video input
- decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)
- Hardware
- 1x RTX PRO 6000 96 GB (measured)
- Quality evidence
- held-out test: planted items found / cited range overlaps the truth shot33 of 33 / 33 of 33 (with the narration; the video survey is the same call)decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route
- Latency
- not measured separately; the survey call is part of the standard run
- Verification
- Proof: strongSelf-host onlyGateway receipt per call on the hosted route.
- In the hosted demo
Standard
narration, scope and quote (hosted demo)
Speech to text for the narration, the video survey, requests and spoken measurements tied to transcript lines, scope lines marked seen, said, seen and said or assumed, a second look at said-only lines, and the priced draft with measure-first and a signed record.
- Models
- Qwen3.8-27B (NVIDIA NVFP4), video input
- MOSS-Transcribe-Diarize 0.9B
- decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)
- Hardware
- 1x RTX PRO 6000 96 GB (measured); the diarizer on a second GPU in the hosted setup
- Quality evidence
- held-out test, 4 walkthroughs: planted items found / lines and conditions that match a planted item33 of 33 / 36 of 37decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route
- held-out test: basis right (seen / said / seen and said) / said-only items wrongly marked seen26 of 33 / 2 of 5decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27
- held-out test: video areas and lengths with the truth inside the range / median error; counts exact2 of 5 / 68.7%; 11 of 14decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27
- all six clips: priced lines whose amount re-computes exactly; spoken quantities used exactly (dev)41 of 41; 3 of 3decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27
- Latency
- measured: about a minute per short walkthrough; a few cents per run at list price
- Verification
- Proof: strongGateway receipt per model call (video calls bind the video's sha256 and sampling); model-call receipt for the speech to text; signed quote record.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.Qwen3.8-27B (NVIDIA NVFP4), video inputnvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance; a long segment is split into sentences with approximate timesMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.9BProof: partialIn the hosted demo | Standard | 0.9B | Proof: partialIn the hosted demo | |
| ||||
Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough) 0 GBProof: partialIn the hosted demo | LiteStandard | 0 GB | Proof: partialIn the hosted demo | |
| ||||
Tools, services and hardware
Tools
- ffmpeg / ffprobe (opens in a new tab)LGPL-2.1+ / GPL-2.0+ (run as separate programs)
Probe the upload, make the clip the model sees, cut keyframes, extract the audio.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:0.1.0GET /walkthrough/info, /walkthrough/samples; POST /walkthrough/draft (SSE or JSON), /walkthrough/runs/{id}/price, /walkthrough/runs/{id}/signoff, /walkthrough/check-arithmetic; GET /walkthrough/runs/{id}/export?format=md|json|csv|record. No GPU; ffmpeg inside. The upload is deleted when the run ends; reports are kept in memory for one hour.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B with image and video input. Internal to the compose network.
- decosa-diarize:8092
MOSS-Transcribe-Diarize for the narration (DECOSA_DIARIZE_URL). No published image yet; built from services/diarize. Without it, send the transcript with the recording.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the hosted Qwen3.8-27B with video runs on one of these cards on our server (KV cache 60.17 GiB after video was turned on), the diarizer on the other.
- 1x RTX 5090 32 GB
Not measured. The weights are about 20 GB; a maximum video request is 32k prompt tokens of KV cache.
- CPU only Does not fit
The steps come from the video model. Clip preparation, keyframes, the record and verification run on CPU.
Latency per lane
- one-minute narrated walkthrough, full draft (speech to text, video survey, narration, scope, second look, pricing), hosted gateway route54.0 s
Measuredmeasured on our server 2026-09-27: 49-58 s on the pre-release server (smoke, rehearsal, recorded runs) and 27-45 s per 33-53 s clip in the eval, while other workloads used the gateway
- a four-minute walkthrough (two video parts)74.0 s
Measuredmeasured on our server 2026-09-27 in the eval (60-94 s across runs)
- re-pricing with the estimator's measurements (code only)50 ms
Measuredmeasured on our server 2026-09-27 (no model call)
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa Walkthrough-to-quote on this machine
You are setting up Walkthrough-to-quote: a phone video of a job walk (rooms or a yard, narrated by the customer or the
estimator) becomes a draft line-item scope and a draft quote. Every line cites a time range and a keyframe and is marked
seen, said or assumed. Quantities from the video are ranges with their method; a dimension nobody said and the video
cannot give stays empty. Prices come only from the business's own price sheet (a CSV), the arithmetic is done in code, and
a record of it is signed by this box's own key. Work step by step, show me each command before you run anything with
`sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/walkthrough-to-quote.zip (2.2 MB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py walkthrough-to-quote` (the api image carries the same bundle under /app/rehearsal/walkthrough-to-quote/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py walkthrough-to-quote --bundle walkthrough-to-quote.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the ceiling water stain (shown, never said) is a seen-only line", "the guest-bath ceiling (said, never filmed) is a said-only line", "the bedroom's length is the one said on camera (14 ft)"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) with video input, and MOSS-Transcribe-Diarize (Apache-2.0) for the narration. The
API is decosa-api (AGPL-3.0-or-later); ffmpeg runs inside it as a separate program (LGPL/GPL).
- Videos of customers' homes stay on this machine. Bind every port to 127.0.0.1. The API deletes each upload when its run
ends and keeps reports in memory for one hour; logs carry counts only. Keep it that way.
- Be honest about what it does: a draft for the estimator. It never prices from anything but the sheet you give it, and
its video quantities are estimates to measure. Film with the customer's knowledge and keep people out of frame.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights; a maximum video request adds
about 32k tokens of KV cache). Measured on an RTX PRO 6000 96 GB. Blackwell cards run NVFP4; on older cards use the
FP8 weights. The diarizer needs about 3 GB more (same card or a second one).
2. `docker --version`, `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them from
the official Docker and NVIDIA repositories after asking me, then check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 35 GB free (model weights, images).
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` and `${DECOSA_REGISTRY}/decosa-llm:<tag>` (**publishing soon**). If a
pull fails, build from source: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a
release that contains `decosa_api/verticals/walkthrough/`, and build `docker/api/Dockerfile` as `decosa-api:local` and
`docker/llm` as `decosa-llm:local` (vLLM 0.29.0, pinned by digest).
- The diarizer has no published image yet: build `services/diarize` from the same checkout as `decosa-diarize:local`
(weights `OpenMOSS-Team/MOSS-Transcribe-Diarize` at revision `704aa4a9c304e8520be88901e0d1960158ef5b15`). Without
it, send the transcript with each video (step 4.4).
- Weights `nvidia/Qwen3.8-27B-NVFP4` at revision `482ca0f3832238542f8f5295dde86b5f22711d80` download on the llm's
first start into the named `hf-cache` volume.
## 3. docker-compose.yml
Write this in `~/decosa/walkthrough/`:
```yaml
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:<tag>
command: ["nvidia/Qwen3.8-27B-NVFP4", "--revision", "482ca0f3832238542f8f5295dde86b5f22711d80",
"--served-model-name", "qwen3.8-27b", "--limit-mm-per-prompt", "{\"image\":4,\"video\":1}",
"--media-io-kwargs", "{\"video\":{\"num_frames\":240,\"fps\":1}}", "--max-model-len", "65536",
"--gpu-memory-utilization", "0.60", "--kv-cache-dtype", "fp8_e4m3", "--host", "0.0.0.0", "--port", "8000"]
ipc: host
volumes: ["hf-cache:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], interval: 15s, retries: 60 }
diarize:
image: decosa-diarize:local
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_DIARIZE_URL: http://diarize:8092
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/walkthrough/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
hf-cache:
```
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not a host folder:
the image runs as uid 10001, and a host folder Docker creates is owned by root (`PermissionError ... /data/keys.sqlite`).
Start with `docker compose up -d`; the llm's first start downloads the weights and takes several minutes.
Video budget: with these flags the model's own processor caps one video at 12,288 video tokens. The hosted service raises
it to 32,768 by mounting a copy of the model's `processor_config.json` and `video_preprocessor_config.json` whose
`video_processor.size.longest_edge` is 67108864 (read-only, over the files in the model folder). Do that only with a
local model folder and room for the extra KV cache; do not use `--mm-processor-kwargs` for it (vLLM 0.29.0 fails at start).
On the first start the api creates this box's Ed25519 key in the volume (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup`, keep it private, never print it.
## 4. Smoke test
1. `curl -s localhost:8445/walkthrough/info | jq '.serving, .steps.asr'`: 1 video, 1 fps, 240 frames; ASR "on".
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"walkthrough-to-quote"}' | jq -r .token)`.
3. Draft the painter sample with its own price sheet:
`curl -s -XPOST localhost:8445/walkthrough/draft -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"painter-bedroom-hall","stream":false}' > draft.json`.
Expect 6-10 lines with a keyframe each; the ceiling water stain as a `seen` line; the guest-bath ceiling as a `said`
line with quantity `missing`; the bedroom's length 14 in `.report.areas[].said_dims`; a subtotal range in
`.report.quote.totals`. On our card it took 40-60 s with other work on the GPU.
4. Your own video without the diarizer: send `"video_b64"` (the file as base64), `"price_sheet_csv"` (your sheet as CSV
text: `code,description,unit,unit_price`; units sq ft, sq yd, lin ft, each, room, day, hour) and
`"transcript": [{"start": 2.1, "end": 3.8, "text": "..."}]` (seconds) instead of `"sample"`. MP4, MOV or WebM, up to
10 minutes and 64 MB.
5. Check the arithmetic and the record: `curl -s localhost:8445/walkthrough/samples/painter-bedroom-hall/prices > prices.csv`,
then `jq -n --slurpfile d draft.json --rawfile p prices.csv '{quote: $d[0].report, price_sheet_csv: $p}' | curl -s -XPOST localhost:8445/walkthrough/check-arithmetic -H 'content-type: application/json' -d @-`
must say `"ok": true`, and `jq '{record: .record}' draft.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
must say `"ok": true`; change one amount in the record and verify again: it must fail.
6. Measure and sign: `curl -s -XPOST localhost:8445/walkthrough/runs/$(jq -r .run_id draft.json)/price -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"measurements":{"1": 250}}'`
re-prices line 1 with your figure (no model call); then `.../signoff` with `{"name":"Your Name","role":"estimator","confirm":true}`.
7. Rehearse on the bundled mock walkthrough before any real one: `python scripts/rehearse.py walkthrough-to-quote --base-url http://127.0.0.1:8445`
(from the decosa-api checkout) must print 11/11 checks passed.
Every model call on the direct route gets a receipt signed with this box's key (status `attested`): an attestation by
me, the operator, not a proof of computation.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /walkthrough/draft` from your
own tooling and keep the Markdown or CSV export and the signed record with each quote. Contract: `API_CONTRACT.md`,
section "walkthrough-to-quote".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds customers' videos.
Follow the provider guide at `/provide` on the site only if I ask, and do not enable it without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
3 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Walkthrough to quote: walk the job with your phone and talk, and the draft quote comes back as line items, each with the moment of the video it comes from, marked seen, said or assumed, with quantities that say where they came from and prices only from your own sheet.
- Who it's for
- Owners and estimators at painting, flooring, remodel, restoration, landscaping and moving businesses who quote from a site visit.
- Where it runs
- Hosted or self-host
- Key numbers
On 4 held-out synthetic walkthroughs it found all 33 planted items and every amount re-computed exactly, but only 2 of 5 areas estimated from the video were within its range, so it asks you to measure first.
- 33 / 33 Planted items found (test split, n = 33)
- 36 / 37 Lines and conditions matching a planted item (test split, n = 37)
- 26 / 33 Basis right (seen / said / seen and said) (test split, n = 33)
- 54.0 s Median end-to-end run, hosted (QA sweep 2026-09-27)
- Models
- Qwen3.8-27B (watches the video, drafts the scope) · MOSS-Transcribe-Diarize (the narration)
- Where
- Hosted or self-host
- Checks
- Receipt per model call, with the video's sha256 and sampling in the request hash; signed ASR receipt; signed record of every line's quantity, unit price and amount
- Industry
- Field and trades
- Input
- Files and media
- Output
- Structured data · Signed record or verdict
- Data
- Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Video understanding · Speaker diarization · Numeric grounding · Signed record
Questions people ask
How does walkthrough to quote turn a video into a quote?
Qwen3.8-27B watches the video at one frame a second and lists the areas and every visible condition with its time; the narration is transcribed and read for requests and spoken measurements; the scope lines cite both. Code works out the quantities and prices them from your sheet.
Where do the prices come from?
Only from the price sheet you give it (a CSV of code, description, unit and unit price). There are no market prices. Amounts are integer cents in code, each unit price is checked against its row, and the arithmetic is signed.
How accurate are the quantities?
Numbers said on camera are used exactly when the transcript has them. Counts from the video were exact 11 of 14 times on held-out clips, but areas and lengths from the video had the truth in range only 2 of 5 times, so each one is marked estimated and the draft lists what to measure first.
What do seen, said and assumed mean?
Seen: the video shows it and nobody asked for it, like a water stain. Said: the narration asks for it but the video does not show it, like a room nobody filmed. Assumed: implied by other work. Each line cites the time and keyframe it comes from.
Can it send the quote to my customer?
No. It is a draft for the estimator: you measure, change or remove lines, re-price and sign, and the business sets the final price. It does not add tax or markup.
Can videos of customers' homes stay on our own hardware?
Yes. The API, Qwen3.8-27B with video input (Apache-2.0) and the speech-to-text model run in containers on one 96 GB GPU, so nothing leaves your machine.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Walkthrough-to-quote
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…