75 · Field and trades · live
Walkthrough-to-quote
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted items found33 / 33test splitn = 33Scope work, damage, hazards and access, as a line or a condition. Dev: 16 / 16.
- Lines and conditions matching a planted item36 / 37test splitn = 37One false condition: broken stair treads on an intact stair.
- Basis right (seen / said / seen and said)26 / 33test splitn = 3330 / 33 if the mover's 'everything in these rooms goes' counts as saying the furniture.
- Said-only items wrongly marked seen2 / 5test splitn = 5
- Cited time range overlaps the truth shot33 / 33test splitn = 33Keyframe inside the shot: 29 / 33
- Video areas and lengths: truth inside the range2 / 5test splitn = 5Median error of the midpoint 68.7%. Dev: 2 / 6, 13.2%.
- Counts from the video exactly right11 / 14test splitn = 14
- Priced amounts that re-compute exactly28 / 28test splitn = 28Integer cents; subtotals 4 / 4
- Seconds per minute of video61.5 stest splitn = 4Median, gateway shared with other workloads
Dataset
6 synthetic narrated job walkthroughs (rendered rooms and a yard, 33-246 s; painting, flooring, remodel, restoration, landscaping, moving) with 49 planted items and ground-truth quantities from the scene geometry; 2 dev (the demo samples), 4 held out.
Caveats
- The same author wrote the scenes, the ground truth and the prompts; the two dev walkthroughs are the demo samples.
- Rendered rooms are cleaner than phone footage and the narration is a clear synthetic voice; real walkthroughs are not measured.
- One run per held-out clip; after the first test run the eval's room matcher was fixed and one dev-motivated change made (visible damage always gets a line); both runs are published.
- A draft for the estimator; it does not measure, and its video areas are often far off.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 54 s
- Receipts
- 5
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.020
Self-host verification
Verified on 27 Sep 2026: Fresh clone of the pre-release branch into a clean directory, docker build of the api image (39 s), the api with a named volume on host networking against the local vLLM (video on) and diarizer, direct route; then torn down.
The rehearsal bundle passed 11 of 11 in 28.6 s; the flooring sample with speech to text inside the container drafted 12 lines (8 priced, $4,037.00-$4,453.50) in 38 s with 4 attested receipts and a model-call receipt for the ASR; the quote record verified; no line, narration or title text in the logs. The local vLLM was the production unit with the 32k video-token override, not the compose default (12,288); the model server's own startup was not re-verified (no new GPU load).
Rehearsal bundle: walkthrough-to-quote.zip (2.2 MB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, live diarizer); production gets this tool when the branch merges.
- Measured on six synthetic walkthroughs of rendered rooms, made and labelled by the building agent; real phone footage is not measured.
- Video area estimates are often far off on held-out clips (2 of 5 inside the range); the draft marks them estimated and ranks them in measure first.
- Runs vary: the demo painter sample came back with 8 to 10 lines across runs; on one run the ceiling stain was a condition only (since then visible damage always gets a line).
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.Qwen3.8-27B (NVIDIA NVFP4), video inputApache-2.0
- The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance; a long segment is split into sentences with approximate timesMOSS-Transcribe-Diarize 0.9BApache-2.0
- Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)AGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · scope and quote from the video only (1)
- held-out test: planted items found / cited range overlaps the truth shot: 33 of 33 / 33 of 33 (with the narration; the video survey is the same call)decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route
Standard · narration, scope and quote (hosted demo) (4)
- held-out test, 4 walkthroughs: planted items found / lines and conditions that match a planted item: 33 of 33 / 36 of 37decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27, gateway route
- held-out test: basis right (seen / said / seen and said) / said-only items wrongly marked seen: 26 of 33 / 2 of 5decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27
- held-out test: video areas and lengths with the truth inside the range / median error; counts exact: 2 of 5 / 68.7%; 11 of 14decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27
- all six clips: priced lines whose amount re-computes exactly; spoken quantities used exactly (dev): 41 of 41; 3 of 3decosa-api docs/evals/walkthrough-to-quote.md, measured on our server 2026-09-27