{"schema_version":"1","site":"https://decosa.ai","id":"walkthrough-to-quote","num":"75","name":"Walkthrough-to-quote","tool_name":"Quote a job from a walkthrough video","short":"Walkthrough quote","blurb":"For painters, flooring, remodel, restoration, landscaping and moving crews. The customer or estimator walks the job with a phone video and talks. It drafts a line-item scope: rooms and areas, quantities with how each was estimated, materials mentioned, and damage, hazards and access problems seen. Every line cites a time range and a keyframe and is marked seen, said or assumed. Quantities from the video are ranges; it never invents a dimension and says what to measure first. Prices come only from your own price sheet, the sums are done in code and signed. A draft for the estimator: you set the final price.","status":"live","labels":{"industry":["field-trades"],"job":["draft","review"],"input":["files"],"deploy":["hosted","selfhost"],"status":"live","output":["data","record"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["field-trades"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["video-understanding","diarize","numeric-grounding","signed-record"],"models":"Qwen3.8-27B (watches the video, drafts the scope) · MOSS-Transcribe-Diarize (the narration)","where":"Hosted or self-host","hardware":"Qwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, pricing and the record on CPU","final_artifact":"A draft quote in Markdown, CSV and JSON: scope lines with time ranges, keyframes, bases, quantities with their method, unit prices and amounts, conditions, where to measure first, and a signed record.","self_host_first":false,"verification":{"receipt_coverage":"partial","summary":"Receipt per model call, with the video's sha256 and sampling in the request hash; signed ASR receipt; signed record of every line's quantity, unit price and amount","manual_qa":{"hosted":{"date":"2026-09-27","result":"pass","p50_ms":54000,"p95_ms":null,"runs":null,"receipts_per_run":5,"cost_per_run_usd":0.0202},"selfhost":{"date":"2026-09-27","result":"pass","method":"Fresh clone of the pre-release branch into a clean directory, docker build of the api image (39 s), the api with a named volume on host networking against the local vLLM (video on) and diarizer, direct route; then torn down.","notes":"The rehearsal bundle passed 11 of 11 in 28.6 s; the flooring sample with speech to text inside the container drafted 12 lines (8 priced, $4,037.00-$4,453.50) in 38 s with 4 attested receipts and a model-call receipt for the ASR; the quote record verified; no line, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load)."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.","Measured on six synthetic walkthroughs of rendered rooms, made and labelled by the building agent; real phone footage is not measured.","Video area estimates are often far off on held-out clips (2 of 5 inside the range); the draft marks them estimated and ranks them in measure first.","Runs vary: the demo painter sample came back with 8 to 10 lines across runs; on one run the ceiling stain was a condition only (since then visible damage always gets a line)."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Planted items found","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Scope work, damage, hazards and access, as a line or a condition. Dev: 16 / 16."},{"name":"Lines and conditions matching a planted item","value":"36 / 37","unit":null,"n":37,"split":"test","note":"One false condition: broken stair treads on an intact stair."},{"name":"Basis right (seen / said / seen and said)","value":"26 / 33","unit":null,"n":33,"split":"test","note":"30 / 33 if the mover's 'everything in these rooms goes' counts as saying the furniture."},{"name":"Said-only items wrongly marked seen","value":"2 / 5","unit":null,"n":5,"split":"test","note":null},{"name":"Cited time range overlaps the truth shot","value":"33 / 33","unit":null,"n":33,"split":"test","note":"Keyframe inside the shot: 29 / 33"},{"name":"Video areas and lengths: truth inside the range","value":"2 / 5","unit":null,"n":5,"split":"test","note":"Median error of the midpoint 68.7%. Dev: 2 / 6, 13.2%."},{"name":"Counts from the video exactly right","value":"11 / 14","unit":null,"n":14,"split":"test","note":null},{"name":"Priced amounts that re-compute exactly","value":"28 / 28","unit":null,"n":28,"split":"test","note":"Integer cents; subtotals 4 / 4"},{"name":"Seconds per minute of video","value":"61.5","unit":"s","n":4,"split":"test","note":"Median, gateway shared with other workloads"}],"dataset":"6 synthetic narrated job walkthroughs (rendered rooms and a yard, 33-246 s; painting, flooring, remodel, restoration, landscaping, moving) with 49 planted items and ground-truth quantities from the scene geometry; 2 dev (the demo samples), 4 held out.","held_out":true,"caveats":["The same author wrote the scenes, the ground truth and the prompts; the two dev walkthroughs are the demo samples.","Rendered rooms are cleaner than phone footage and the narration is a clear synthetic voice; real walkthroughs are not measured.","One run per held-out clip; after the first test run the eval's room matcher was fixed and one dev-motivated change made (visible damage always gets a line); both runs are published.","A draft for the estimator; it does not measure, and its video areas are often far off."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/walkthrough-to-quote"},"stack":{"summary":"The customer or estimator walks the job with a phone and talks. Qwen3.8-27B watches the video and lists the areas and everything visible that matters (damage, hazards, access, fixtures), each with its time; the narration is transcribed and read for requests and spoken measurements; the scope lines cite both and are marked seen, said, seen and said, or assumed. Code works out quantities (said numbers only if the transcript has them, video estimates as ranges with their method, otherwise a request to measure), prices them from your sheet in integer cents and signs the arithmetic. The estimator measures, re-prices and signs; the business sets the final price. For painters, flooring, remodel, restoration, landscaping and moving crews.","tagline":"A narrated phone walkthrough of a job, turned into cited scope lines and a draft quote priced only from your own price sheet.","deployment":"hosted-or-self-host","regulatory_note":"Not legal advice; a draft for the estimator, not a contract or a sendable quote. Checked 27 Sep 2026: if the quote leads to a sale agreed at the customer's home, the FTC's Cooling-Off Rule may apply: 16 CFR 429.0(a) defines a door-to-door sale as one where the buyer's agreement is made 'at a place other than the place of business of the seller', at $25 or more at the buyer's residence, and the buyer then has three business days to cancel (https://www.ecfr.gov/current/title-16/chapter-I/subchapter-D/part-429, read via the Cornell LII copy at https://www.law.cornell.edu/cfr/text/16/429.0). Many states also set rules for home-improvement contracts; we did not check them. Recording a conversation: in some US states everyone must consent (California Penal Code 632(a) makes it an offence to record a confidential communication 'without the consent of all parties', https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?sectionNum=632.&lawCode=PEN); film with the customer's knowledge. Video of a home can be personal data; keep people out of frame. The demo walkthroughs are synthetic.","components":[{"id":"llm","role":"Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.","name":"Qwen3.8-27B (NVIDIA NVFP4), video input","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 with --limit-mm-per-prompt image=4,video=1 and --media-io-kwargs fps=1, 240 frames","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"asr","role":"The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance; a long segment is split into sentences with approximate times","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package (decosa-api services/diarize)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"quote","role":"Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)","name":"decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; ffmpeg and ffprobe","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · scope and quote from the video only","summary":"The areas, conditions and scope lines the video shows, each with a time range and a keyframe, priced from your sheet. No speech to text: send a transcript if you have one, otherwise nothing is marked said and spoken measurements are not used.","components":["llm","quote"],"hardware":"1x RTX PRO 6000 96 GB (measured)","quality_evidence":[{"metric":"held-out test: planted items found / cited range overlaps the truth shot","value":"33 of 33 / 33 of 33 (with the narration; the video survey is the same call)","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route"}],"latency_note":"not measured separately; the survey call is part of the standard run","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":"Gateway receipt per call on the hosted route.","hosting":null},{"id":"standard","label":"Standard · narration, scope and quote (hosted demo)","summary":"Speech to text for the narration, the video survey, requests and spoken measurements tied to transcript lines, scope lines marked seen, said, seen and said or assumed, a second look at said-only lines, and the priced draft with measure-first and a signed record.","components":["llm","asr","quote"],"hardware":"1x RTX PRO 6000 96 GB (measured); the diarizer on a second GPU in the hosted setup","quality_evidence":[{"metric":"held-out test, 4 walkthroughs: planted items found / lines and conditions that match a planted item","value":"33 of 33 / 36 of 37","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route"},{"metric":"held-out test: basis right (seen / said / seen and said) / said-only items wrongly marked seen","value":"26 of 33 / 2 of 5","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"},{"metric":"held-out test: video areas and lengths with the truth inside the range / median error; counts exact","value":"2 of 5 / 68.7%; 11 of 14","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"},{"metric":"all six clips: priced lines whose amount re-computes exactly; spoken quantities used exactly (dev)","value":"41 of 41; 3 of 3","source":"decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27"}],"latency_note":"measured: about a minute per short walkthrough; a few cents per run at list price","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Gateway receipt per model call (video calls bind the video's sha256 and sampling); model-call receipt for the speech to text; signed quote record.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"GET /walkthrough/info, /walkthrough/samples; POST /walkthrough/draft (SSE or JSON), /walkthrough/runs/{id}/price, /walkthrough/runs/{id}/signoff, /walkthrough/check-arithmetic; GET /walkthrough/runs/{id}/export?format=md|json|csv|record. No GPU; ffmpeg inside. The upload is deleted when the run ends; reports are kept in memory for one hour."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B with image and video input. Internal to the compose network."},{"name":"decosa-diarize","port":8092,"image":null,"purpose":"MOSS-Transcribe-Diarize for the narration (DECOSA_DIARIZE_URL). No published image yet; built from services/diarize. Without it, send the transcript with the recording."}],"tools":[{"name":"ffmpeg / ffprobe","url":"https://ffmpeg.org","license":"LGPL-2.1+ / GPL-2.0+ (run as separate programs)","purpose":"Probe the upload, make the clip the model sees, cut keyframes, extract the audio."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured: the hosted Qwen3.8-27B with video runs on one of these cards on our server (KV cache 60.17 GiB after video was turned on), the diarizer on the other."},{"tier":"1x RTX 5090 32 GB","fits":null,"notes":"Not measured. The weights are about 20 GB; a maximum video request is 32k prompt tokens of KV cache."},{"tier":"CPU only","fits":false,"notes":"The steps come from the video model. Clip preparation, keyframes, the record and verification run on CPU."}],"latency":[{"lane":"one-minute narrated walkthrough, full draft (speech to text, video survey, narration, scope, second look, pricing), hosted gateway route","typical_ms":54000,"source":"measured on our server 2026-09-27: 49-58 s on the pre-release server (smoke, rehearsal, recorded runs) and 27-45 s per 33-53 s clip in the eval, while other workloads used the gateway"},{"lane":"a four-minute walkthrough (two video parts)","typical_ms":74000,"source":"measured on our server 2026-09-27 in the eval (60-94 s across runs)"},{"lane":"re-pricing with the estimator's measurements (code only)","typical_ms":50,"source":"measured on our server 2026-09-27 (no model call)"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"What it does","value":"Turns one narrated walkthrough video and your price sheet into scope lines (each with a time range, a keyframe and a basis: seen, said, seen and said, or assumed), conditions (damage, hazards, access), materials mentioned, quantities with where they came from, prices from your sheet, a subtotal range and the lines to measure first. You enter measurements, re-price, and sign; exports in Markdown, CSV and JSON."},{"label":"What it does not do","value":"It does not price from anything but your sheet (no market prices), add tax, markup or minimum charges, or send a quote to a customer. It does not see behind walls or under floors, check codes or permits, or measure: video quantities are estimates, and on held-out clips areas were often far off. Videos over 4 minutes are watched in 2-minute parts at one frame a second, so brief views can be missed."},{"label":"Data retention","value":"The video and its clip are deleted when the run ends. The report, with its keyframes, is kept in memory for one hour for the token or key that made it. Logs carry run ids and counts, never titles, narration or line text. You keep the exports and the signed record, which holds hashes and numbers only."},{"label":"What leaves the box (hosted demo)","value":"The clip (no audio), the narration text and your price sheet's codes and descriptions go to Qwen3.8-27B through our gateway, which Decosa operates; the audio is transcribed on Decosa's hosted service. The gateway's receipts hold hashes, not content. Self-hosted, nothing leaves."},{"label":"Accuracy","value":"On 4 held-out synthetic walkthroughs: 33 of 33 planted items found, 36 of 37 lines and conditions matched a planted item, the basis right on 26 of 33 (2 of 5 said-only items were wrongly marked seen), counts exact on 11 of 14, but only 2 of 5 video areas and lengths had the truth inside the range (median error 68.7%). Every priced amount re-computed exactly."},{"label":"Cost per walkthrough","value":"A few cents for a short walkthrough at gateway list price (a few model calls, measured on the demo sample), plus speech to text on Decosa's hosted service."},{"label":"Output","value":"A draft quote with lines, time ranges, keyframes, bases, quantities and their method, unit prices and amounts; a decosa.record.v1 record of the arithmetic, re-signed when you re-price and when you sign (verify at /record/verify), and POST /walkthrough/check-arithmetic to re-run the sums."}],"hosted_now":{"needs":["diarize","qwen3.8-27b","video-input"],"off":["diarize","video-input"],"live_by_default":false,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#walkthrough-to-quote","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"The video and its clip are deleted when the run ends. The report, with its keyframes, is kept in memory for one hour for the token or key that made it. Logs carry run ids and counts, never titles, narration or line text. You keep the exports and the signed record, which holds hashes and numbers only.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/operations/walkthrough-to-quote","input":"walkthrough","lanes":[{"id":"survey","title":"What the video shows","kind":"list"},{"id":"scope","title":"Scope lines","kind":"list"},{"id":"quote","title":"Priced from your sheet","kind":"list"},{"id":"record","title":"Quote record","kind":"json"}],"samples":[{"n":1,"id":"painter-bedroom-hall","title":"Painter bedroom hall","deep_link":"/tools/operations/walkthrough-to-quote?sample=1&autorun=0"},{"n":2,"id":"flooring-living-dining","title":"Flooring living dining","deep_link":"/tools/operations/walkthrough-to-quote?sample=2&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/walkthrough-to-quote-hosted.md","selfhost":"/prompts/walkthrough-to-quote-selfhost.md","assemble":"/prompts/walkthrough-to-quote-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/walkthrough-to-quote.zip","bundle_url":"https://decosa.ai/samples/walkthrough-to-quote.zip","folder":"/samples/walkthrough-to-quote/","expected":"/samples/walkthrough-to-quote/expected.json","files":["/samples/walkthrough-to-quote/expected.json","/samples/walkthrough-to-quote/inputs/prices.csv","/samples/walkthrough-to-quote/inputs/transcript.json","/samples/walkthrough-to-quote/inputs/walkthrough.mp4"],"bytes":2346108,"checks":["the ceiling water stain (shown, never said) is a seen-only line","the guest-bath ceiling (said, never filmed) is a said-only line","the bedroom's length is the one said on camera (14 ft)","every line has a keyframe","at least five scope lines","every quantity is said, estimated with a method, per room or missing: none is made up","the arithmetic re-computes exactly from the price sheet","a price sheet without the required columns is refused","the signed record of the arithmetic verifies","the record fails once it is changed","every model call has a signed receipt"],"licence":"Synthetic: rooms rendered with three.js (MIT) from a scripted scene written for Decosa, narrated by the Kokoro-82M stock voice am_michael (Apache-2.0). No real homes, people or voices. Video, transcript and price sheet CC0; part of decosa-api, which will be released under AGPL-3.0-or-later; until then the source is on request.","about":"A 53-second walkthrough of a rendered bedroom and hallway, narrated by a stock synthetic voice. The narrator gives the bedroom's size (twelve by fourteen, eight-foot ceiling), asks for the walls, the trim and both closet doors, points at peeling paint, never mentions a water stain the camera shows on the ceiling, and asks for a guest-bath ceiling that is never filmed. The draft must list the stain as a seen-only line and the guest-bath ceiling as a said-only line, use the size said on camera for the bedroom, cite a keyframe for every line, price only from the painter's sheet with arithmetic that re-computes exactly, and seal a record that verifies and fails once changed.","run":{"containers":"docker compose exec api python scripts/rehearse.py walkthrough-to-quote","checkout":"python scripts/rehearse.py walkthrough-to-quote --bundle walkthrough-to-quote.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py walkthrough-to-quote"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=walkthrough-to-quote","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":61.6,"basis":"estimate","unknown":[]}],"mac":null},"links":{"page":"/tools/operations/walkthrough-to-quote","json":"/use-cases/walkthrough-to-quote.json","metrics":"/metrics/walkthrough-to-quote","console":"/tools/operations/walkthrough-to-quote","console_sample":"/tools/operations/walkthrough-to-quote?sample=1&autorun=0","stack":"/tools/operations/walkthrough-to-quote#stack","try_live":"/tools/operations/walkthrough-to-quote","watch":"/tools/operations/walkthrough-to-quote","build":"/tools/operations/walkthrough-to-quote#build","self_host":"/tools/operations/walkthrough-to-quote#self-host","prompts":{"hosted":"/prompts/walkthrough-to-quote-hosted.md","selfhost":"/prompts/walkthrough-to-quote-selfhost.md","assemble":"/prompts/walkthrough-to-quote-assemble.md","mac":null}}}