Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Walkthrough-to-quote eval (27 Sep 2026)

Six synthetic job walkthroughs run through the real pipeline on the metered gateway route (Qwen3.8-27B with video input; the narration from MOSS-Transcribe-Diarize, errors included) with each clip's own price sheet. Script: scripts/eval_walkthrough_to_quote.py; clips built by scripts/walkthrough_clips/ (scene specs with the ground truth in scenes.py, the three.js renderer, the Kokoro narration and the muxer). Results: docs/evals/walkthrough-to-quote/ (results-v2.json is the shipped pipeline; results-v1-test*.json the first test run; one run-*.json per clip and run).

Data

  • Walkthroughs: rooms and a yard with known dimensions, rendered frame by frame on CPU (three.js r171 on SwiftShader in headless Chromium, 10 fps, 1280x720), with a hand-held camera path; cuts between rooms. Fictional homes, no people. Six trades: a painter (bedroom and hall, 53 s), flooring (living, dining, entry, 41 s), a bathroom remodel (33 s), a flooded basement (41 s), a backyard fence and lawn (40 s) and a moving inventory (three rooms and a stair, 246 s, the long clip).
  • Narration: Kokoro-82M stock voices am_michael and af_heart (Apache-2.0; Decosa's house voices, synthetic, not a real person's), placed at each shot's start, then transcribed by the diarizer like any upload. The eval uses the diarizer's transcript, errors included ("Bassboards" for baseboards; the basement narration came back as one 38 s segment, which the pipeline splits into sentences with approximate times).
  • Planted in every walkthrough (49 items): scope work that is seen and said, damage or hazards that are seen but never said (a ceiling water stain, cracked tiles, mold, an open junction box, a wasp nest, bare lawn), and work that is said but never shown (a guest-bath ceiling, 13 stairs, an exhaust fan, a neighbour's oak, a chest freezer, drying days). Ground-truth quantities come from the scene geometry (floor and net wall areas, baseboard runs, fence length, counts). Each item's time truth is the shot where it is shown or said.
  • Price sheets: one fictional sheet per trade (6-8 rows each, scenes.py).
  • Split: dev = painter-bedroom-hall, flooring-living-dining (the two demo samples; used while writing the prompts). Test = the other four, run with frozen prompts.
  • Licence: everything is synthetic and ours (CC0 renders; the generator is AGPL-3.0-or-later with decosa-api; three.js MIT).

Metrics

A planted item is found when a line or a condition has one of its keywords, none of its excluded words, and its area (named in the text, or the cited time falls in shots filmed in that room). One candidate per item, the most constrained item first; a second line about a found item is a duplicate, not an error. Precision counts every line and condition not marked assumed.

Metric Dev (2 clips, 16 items) Test (4 clips, 33 items)
Planted items found (recall) 16 / 16 33 / 33
Lines and conditions that match a planted item (precision) 16 / 16 36 / 37
Basis right (seen / said / seen and said) 14 / 16 26 / 33
Said-only items wrongly marked seen 0 / 2 2 / 5
Seen-only items marked said as well 0 / 2 5 / 11 (4 are the mover's "everything in these rooms goes", see below)
Cited time range overlaps the truth shot (±3 s) 16 / 16 33 / 33
Keyframe inside the truth shot (±3 s) 12 / 16 29 / 33
Spoken quantities used exactly 3 / 3 no spoken quantity planted on test
Areas and lengths from the video: truth inside the range 2 / 6 2 / 5
Areas and lengths from the video: median error of the range's midpoint 13.2% 68.7% (0, 11.4, 68.7, 248.9, 700%)
Counts from the video exactly right 1 / 3 11 / 14
Priced lines whose amount re-computes exactly (integer cents) 13 / 13 28 / 28
Quotes whose subtotals re-compute exactly / recompute ok 2 / 2 4 / 4
Invented dimensions (a spoken number not in the transcript) kept 0 0 (1 dropped by the guard)
Model calls, all with a gateway-signed receipt 8 17
Seconds per minute of video (gateway shared with other workloads) 38 median 61.5 median (18 on the 4-minute clip)

What went wrong on the test split

  • Area estimates from video are the weak part. The model underestimated the basement floor (119-182 sq ft for 480), sized the tub surround as a whole wall (133-181 sq ft for 45) and the bare lawn patches as the lawn (1,200-2,000 sq ft for 200). Only 2 of 5 held-out video areas and lengths had the truth inside the range. The draft marks every such quantity "estimated from the video" with its method, and "measure first" ranks them by the money they move, but a quote priced from them without a measurement would be badly wrong. This is why it is a draft for the estimator.
  • Counts are better but not exact: 2 broken fence boards for 5, 4 cracked bathroom tiles for 2, 20 moving boxes for 18 (stacked boxes are hard at one frame a second).
  • Two said-only items were marked seen. "Probably three days with fans and dehumidifiers" was tied to the flood-line items, so the drying line came back seen and said, with the flood line's length as its quantity; the price sheet prices drying per day, so the unit check refused to price it (unit mismatch) rather than quote it. "The floor feels soft by the toilet" was tied to the cracked tiles. The exhaust fan (never installed, so never visible) stayed said only.
  • Four seen-only moving items came back seen and said: the narrator opens with "everything in these three rooms is going", and the model tied the armchair, TV stand, dresser and bookshelf to it. That reading is defensible; the planted label is stricter. Counting those four as right, the basis is right on 30 of 33.
  • One false condition: "broken or missing stair treads" on the mover's stair, which is intact (the render's tread edges may read as gaps). Counted against us.
  • Missing quantities where a count was easy: the gate and the dead shrub (one each) came back "needs a count" because the survey gave no count; the draft asks rather than guesses.
  • Keyframes: 4 of 33 keyframes fall outside the shot that shows the item (the ranges the model cited are wide; the keyframe is its "clearest" moment, which it sometimes places at the wrong end).

Long clip: chunked or not (ablation)

The pipeline surveys videos over 150 s in 120 s chunks (10 s overlap) and adds each chunk's offset in code, because the video-input eval saw the clock stretch on long clips. On the one long clip (246 s), run once each way: chunked 11 / 11 found, 11 / 11 keyframes in the truth shot, 5 calls, 60 s; one call 11 / 11 and 11 / 11, 4 calls, 59 s, with two ranges drawn wider (2:55-4:05). No drift showed either way on this clip, so this does not prove chunking helps; it is kept as the guard the video-input eval motivates, at the cost of one call per extra chunk.

Changes made on dev (before the test split ran)

  1. The survey and scope prompts: size walls by counting references (door widths, tiles, planks, fence sections); name the specific thing and place in each line; for new work, cite the thing it replaces or the place it goes; every request must be a line. The said-only check is told the room in brackets must be the one on screen.
  2. A survey reply cut off by the token limit is repaired (closed after its last complete element) and an unreadable one is asked once more (dev1 lost a whole flooring survey to a bad reply).
  3. A request no line cites is added as a said-only line; a hazard nothing cites is added as a condition.

After the first test run (results-v1-test.json), two changes, neither tuned on test data: the eval's area matcher now also accepts a line whose time falls in the room's shots (the basement came back as "Main room" and 6 of 7 found items scored as missed; results-v1-test-rescored.json re-scores the saved v1 answers: 33 / 33 found), and visible damage that no line cites becomes a seen-only line even when a condition mentions it (a smoke run on the dev painter sample lost the ceiling stain to a condition). results-v2.json is the shipped pipeline on all six clips; the v1 test numbers were: 33 / 33 found, 36 / 37 precision, basis 25 / 33, arithmetic 29 / 29 exact.

Expected properties of the demo sample (for the rehearsal bundle)

rehearsal/walkthrough-to-quote (the painter walk with its transcript and price sheet) checks, against any install:

  1. the ceiling water stain (shown, never said) is a seen-only line;
  2. the guest-bath ceiling (said, never filmed) is a said-only line;
  3. the bedroom's length is the 14 ft said on camera;
  4. every line has a keyframe, and there are at least five lines;
  5. every quantity is said, estimated, per room or missing (none is made up), and the arithmetic re-computes exactly;
  6. a sheet without the required columns is refused; the signed record verifies and fails once changed; every model call has a signed receipt. Passed 11 / 11 against the pre-release server (51 s). The smoke module (scripts/smoke/walkthrough-to-quote.py) runs the same sample: 4 calls, 54,154 tokens, about $0.020 at list price, 58 s.

Limits

  • Six clips made and labelled by the building agent; rendered rooms are cleaner than phone footage (no motion blur, clutter, poor light, fingers over the lens), and the narration is a clear synthetic voice. Real walkthroughs are not measured.
  • One run per clip on test; temperature 0 is not deterministic under batching (the dev painter sample came back with 8 to 10 lines across runs).
  • Quantities: video area estimates are poor on held-out clips; counts are fair; spoken numbers are exact when the transcript has them. Roofs, HVAC equipment and exterior siding are not in the set.
  • The price sheet has no markup, tax, minimum-charge or percentage rows; the quote is a subtotal before those.
  • A draft for the estimator. It does not check codes, permits or safety, and does not see hidden conditions.