Skip to content
decosa

05 · Field and trades · live

Field reports

Open the toolJSON

Eval results

Not held outRun 23 Sep 2026

  • Expected issues in the report (grounded report + safety sweep + claim check)40/41 (98%); before 38/41 (93%)syntheticn = 41
  • Severity within the accepted range, of issues found40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)syntheticn = 40
  • Stated facts kept51/51; before 32/51 (63%)syntheticn = 51incl. passed checks and spec limits
  • Claims the judge flagged as contradicted or unsupported5/125syntheticn = 1252 are the overall-condition rating; 3 are speech-recognition spellings quoted verbatim. By hand: no invented findings.

Dataset

Internal synthetic eval: 8 TTS inspection walk-throughs (2 recorded demo + 6 new) with a gold checklist and a Gemma 4 judge, run through the live stack on decosa-api b24ab2a. The TTS audio was macOS system voices; this eval was not re-run after the demo audio was re-voiced with Decosa house voices (Kokoro-82M) on 26 Sep 2026.

Caveats

  • The prompts were tuned on these same 8 scripts, so there is no held-out set.
  • Synthetic TTS audio only; no real site recordings.
  • Small n (8 walk-throughs).
  • Model judge (Gemma 4), checked by hand only in part.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
26 s
Receipts
28
Model calls
n/a
Tokens
n/a
Cost per run
$0.026

Self-host verification

Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Rehearsal bundle: field.zip (665 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • The report is written after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load.
  • Measurements are checked against the spec the inspector states; there is no built-in code database.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · 4-bit on a 32 GB card, plus a small card for speech (3)
  • Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM: 478coding-agent-bench README (our server)
  • Same weights served by llama.cpp (Q4 GGUF path): avoid for this tier: 62–63 / 500coding-agent-bench README (our server)
  • Field lane quality: not measured yet
Standard · one 96 GB card (the hosted demo) (8)
  • Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM: 478coding-agent-bench README (our server)
  • ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; FP8 on vLLM): 34.2scribe-bench wiki/models.md
  • Quality smoke, math / code: 38/40, 12/12Decosa model benchmarks (Sep 2026)
  • Field report, internal synthetic eval (n = 8 walk-throughs: 2 recorded demo + 6 new), with the grounded report, safety sweep and claim check: expected issues in the report: 40/41 (98%); before 38/41 (93%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
  • Same eval: severity within the accepted range, of issues found: 40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
  • Same eval: stated facts kept (incl. passed checks and spec limits, now fields of their own): 51/51; before 32/51 (63%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
  • Same eval: missed safety item (bathroom receptacles without GFCI): in the report in 3 of 3 re-runs. Run on the original report, the safety sweep adds it as critical, quoted at 1:03eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
  • Same eval: claims the judge flagged as contradicted or unsupported: 5/125 (before 0/100, plus 1 found by hand: 'cracked chimney' where the inspector said sealant). By hand: no invented findings. 2 of the 5 are the report's overall-condition rating; 3 are speech-recognition spellings now quoted verbatim ('ceiling' for sealant, 'pole' for pull stations, 'stab block' for Stab-Lok), which need an ASR glossary, not a model changeeval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
Best · two 96 GB cards (4)
  • ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; run on a Mac via MLX, not the NVFP4 kit): 35.8scribe-bench wiki/models.md
  • Medical-term miss rate on PriMock57, MOSS-Transcribe-Diarize vs Nemotron-3.5 streaming (%): 8.4 vs 12.7scribe-bench RESULTS.md / wiki/decoder-finding.md
  • Owner's coding benchmark, core tasks (/500), DeepSeek V4-Flash: 453–459coding-agent-bench README (our server)
  • Field lane quality: not measured yet
Wanted · the largest open flash models (1)
  • Field lane quality: not measured yet

How we measure · All tools