{"schema_version":"1","site":"https://decosa.ai","id":"field","num":"05","name":"Field reports","tool_name":"Write the inspection report","short":"Field","blurb":"Talk through an inspection or site walk. Get the checklist, issues by severity, measurements and a finished report.","status":"live","labels":{"industry":["field-trades"],"job":["transcribe","draft"],"input":["voice"],"deploy":["hosted","selfhost"],"status":"live","output":["text","data"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["field-trades"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["live-asr"],"models":"Voxtral 4B · Qwen3.8-27B","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB), or 2× RTX 5090","final_artifact":"A structured report (JSON) plus markdown.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Strong on text lanes","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":26000,"p95_ms":null,"runs":null,"receipts_per_run":28,"cost_per_run_usd":0.026},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers","notes":"The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified."},"known_limits":["The report is written after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load.","Measurements are checked against the spec the inspector states; there is no built-in code database."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Expected issues in the report (grounded report + safety sweep + claim check)","value":"40/41 (98%); before 38/41 (93%)","unit":null,"n":41,"split":"synthetic","note":null},{"name":"Severity within the accepted range, of issues found","value":"40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)","unit":null,"n":40,"split":"synthetic","note":null},{"name":"Stated facts kept","value":"51/51; before 32/51 (63%)","unit":null,"n":51,"split":"synthetic","note":"incl. passed checks and spec limits"},{"name":"Claims the judge flagged as contradicted or unsupported","value":"5/125","unit":null,"n":125,"split":"synthetic","note":"2 are the overall-condition rating; 3 are speech-recognition spellings quoted verbatim. By hand: no invented findings."}],"dataset":"Internal synthetic eval: 8 TTS inspection walk-throughs (2 recorded demo + 6 new) with a gold checklist and a Gemma 4 judge, run through the live stack on decosa-api b24ab2a. The TTS audio was macOS system voices; this eval was not re-run after the demo audio was re-voiced with Decosa house voices (Kokoro-82M) on 26 Sep 2026.","held_out":false,"caveats":["The prompts were tuned on these same 8 scripts, so there is no held-out set.","Synthetic TTS audio only; no real site recordings.","Small n (8 walk-throughs).","Model judge (Gemma 4), checked by hand only in part."],"date":"2026-09-23","doc_url":null},"stack":{"summary":"The inspector narrates the walk-through into a phone or laptop. Speech is transcribed live, and every ~10 seconds the language model updates a checklist for the inspection type, an issue list graded critical, major, minor or info, the measurements heard with any spec the inspector states, and the checks that passed. Each item carries the inspector's words, quoted verbatim from the transcript, with the time they were said. On stop it writes a structured report (JSON plus markdown). A safety sweep then checks the transcript against common safety items for the inspection type and adds any the inspector mentioned but the report left out. Finally, a claim check tests every report item against the transcript lines it cites: unsupported items are removed and listed for review, overstated ones are corrected. The report is ready to export as JSON or PDF. Self-hosted, the whole stack runs on one GPU box on-site and keeps working without an internet connection once the weights are downloaded.","tagline":"Talk through an inspection, site visit or job; get the checklist, issues by severity, measurements and a finished report.","deployment":"hosted-or-self-host","regulatory_note":"Drafting aid: the report is a draft for the inspector to review and sign; it does not replace a licensed inspection or code-compliance determination. Self-host keeps site audio, transcripts and reports on your hardware.","components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602","license":"Apache-2.0","params":"4.43B","quant":"BF16 (8.3 GB weights; compose allots 0.25 of a 96 GB card)","vram_gb":24,"memory_gb_estimate":null,"engine":"vLLM realtime WebSocket (/v1/realtime)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":[],"alternative_to":null},{"id":"llm","role":"Lanes and report (language model)","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 MLP + FP8 attention/GDN, FP8 KV cache (19.9 GiB weights measured; compose allots 0.60 of a 96 GB card). Hopper/Ada: Qwen/Qwen3.8-27B-FP8.","vram_gb":57.6,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, MTP speculative decoding (3 draft tokens), --language-model-only","receipt_coverage":"strong","in_hosted_demo":null,"tiers":[],"alternative_to":null},{"id":"asr_final","role":"Final transcript pass (speaker-attributed)","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.91B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"Hugging Face transformers (trust_remote_code), offline pass over the recorded walk-through","receipt_coverage":"none","in_hosted_demo":null,"tiers":["best"],"alternative_to":null},{"id":"llm_best","role":"Lanes and report (larger model)","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts + FP8 (about 159–176 GB of weights)","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X community build, TP2 on 2× RTX PRO 6000, MTP draft fixed by the kit's patches","receipt_coverage":"none","in_hosted_demo":null,"tiers":["best"],"alternative_to":null},{"id":"llm_wanted","role":"Lanes and report (network-hosted flash model)","name":"GLM-5.3-Flash or DeepSeek-V4.1-Flash","hf_repo":"zai-org/GLM-5.3-Flash | deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"321B (GLM-5.3-Flash) / 763B incl. Engram tables (V4.1-Flash)","quant":null,"vram_gb":null,"memory_gb_estimate":170,"engine":null,"receipt_coverage":"none","in_hosted_demo":null,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · 4-bit on a 32 GB card, plus a small card for speech","summary":"Same models and weights as Standard, so same output quality; you give up context length (16k) and concurrency (1–2 walk-throughs at a time).","components":["asr","llm"],"hardware":"1× 32 GB Blackwell card (NVFP4 LLM) + 1× ≥16 GB card (Voxtral)","quality_evidence":[{"metric":"Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM","value":"478","source":"coding-agent-bench README (our server)"},{"metric":"Same weights served by llama.cpp (Q4 GGUF path): avoid for this tier","value":"62–63 / 500","source":"coding-agent-bench README (our server)"},{"metric":"Field lane quality","value":"not measured yet","source":null}],"latency_note":"estimate: similar per-call latency to Standard at 1 session; not measured on a 32 GB card","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":"Qwen3.8-27B is the hosted model qwen3.8-27b: gateway-signed receipts on the gateway route; the default self-host route (direct) signs receipts with the box's own key (attested).","hosting":null},{"id":"standard","label":"Standard · one 96 GB card (the hosted demo)","summary":"Live Voxtral transcript plus Qwen3.8-27B NVFP4 with MTP; about 4 walk-throughs at once, 64k context.","components":["asr","llm"],"hardware":"1× RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM","value":"478","source":"coding-agent-bench README (our server)"},{"metric":"ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; FP8 on vLLM)","value":"34.2","source":"scribe-bench wiki/models.md"},{"metric":"Quality smoke, math / code","value":"38/40, 12/12","source":"Decosa model benchmarks (Sep 2026)"},{"metric":"Field report, internal synthetic eval (n = 8 walk-throughs: 2 recorded demo + 6 new), with the grounded report, safety sweep and claim check: expected issues in the report","value":"40/41 (98%); before 38/41 (93%)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: severity within the accepted range, of issues found","value":"40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: stated facts kept (incl. passed checks and spec limits, now fields of their own)","value":"51/51; before 32/51 (63%)","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: missed safety item (bathroom receptacles without GFCI)","value":"in the report in 3 of 3 re-runs. Run on the original report, the safety sweep adds it as critical, quoted at 1:03","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"},{"metric":"Same eval: claims the judge flagged as contradicted or unsupported","value":"5/125 (before 0/100, plus 1 found by hand: 'cracked chimney' where the inspector said sealant). By hand: no invented findings. 2 of the 5 are the report's overall-condition rating; 3 are speech-recognition spellings now quoted verbatim ('ceiling' for sealant, 'pole' for pull stations, 'stab block' for Stab-Lok), which need an ASR glossary, not a model change","source":"eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)"}],"latency_note":"measured on our server: lanes in seconds; on stop, the report, safety sweep and claim check (direct route) bring the finished report seconds after the audio ends","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Qwen3.8-27B is the hosted model qwen3.8-27b: gateway-signed receipts on the gateway route; the default self-host route (direct) signs receipts with the box's own key (attested).","hosting":null},{"id":"best","label":"Best · two 96 GB cards","summary":"DeepSeek-V4-Flash writes the lanes and report; a final MOSS-Transcribe-Diarize pass cleans the transcript. Needs the whole two-card box and a community vLLM build.","components":["asr","asr_final","llm_best"],"hardware":"2× RTX PRO 6000 Blackwell 96 GB (TP2)","quality_evidence":[{"metric":"ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; run on a Mac via MLX, not the NVFP4 kit)","value":"35.8","source":"scribe-bench wiki/models.md"},{"metric":"Medical-term miss rate on PriMock57, MOSS-Transcribe-Diarize vs Nemotron-3.5 streaming (%)","value":"8.4 vs 12.7","source":"scribe-bench RESULTS.md / wiki/decoder-finding.md"},{"metric":"Owner's coding benchmark, core tasks (/500), DeepSeek V4-Flash","value":"453–459","source":"coding-agent-bench README (our server)"},{"metric":"Field lane quality","value":"not measured yet","source":null}],"latency_note":"measured single-stream decode 108.8 → 150.6 tok/s with MTP (dsv4-flash-nvfp4-sm120 README); field lane latency not measured","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"No receipts today: neither DeepSeek-V4-Flash nor the MOSS pass is a hosted model. Only the live Voxtral step is shared with Standard.","hosting":null},{"id":"wanted","label":"Wanted · the largest open flash models","summary":"GLM-5.3-Flash or DeepSeek-V4.1-Flash writes the report, served by network providers; speech recognition stays on your machine. Not served yet.","components":["asr","llm_wanted"],"hardware":"Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). Speech recognition stays local. Estimate.","quality_evidence":[{"metric":"Field lane quality","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Field pack lane engine: /ws/live?vertical=field, /demo/replay, /receipts/{id}, /healthz. No GPU. Publishing soon; builds from docker/api."},{"name":"llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"Qwen3.8-27B on vLLM, OpenAI-compatible, served as qwen3.8-27b. Internal to the compose network."},{"name":"asr","port":8000,"image":"${DECOSA_REGISTRY}/decosa-asr:0.1.0","purpose":"Voxtral realtime WebSocket, served as voxtral-realtime. Internal to the compose network."}],"tools":[],"hardware":[{"tier":"1× RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Compose defaults (LLM 0.60, ASR 0.25 of the card), about 4 simultaneous walk-throughs. The latencies below were measured on our server with the two models on separate cards of this type."},{"tier":"1× H100 / H200 (80–141 GB)","fits":true,"notes":"No NVFP4 on Hopper: LLM_MODEL=Qwen/Qwen3.8-27B-FP8, LLM_GPU_UTIL=0.62. Starting point, not measured."},{"tier":"1× L40S / RTX 6000 Ada 48 GB","fits":true,"notes":"FP8 checkpoint, LLM_MAX_LEN=16384, LLM_GPU_UTIL=0.70, ASR_GPU_UTIL=0.22, DECOSA_LIVE_CAP=2. A walk-through uses well under 16k tokens. Not measured."},{"tier":"1× 24–32 GB card","fits":false,"notes":"A 4-bit Qwen3.8-27B fits on its own (NVFP4 weights 19.9 GiB; Q4 GGUF 17.5–21 GB), but not together with the ASR model and a useful KV cache on the same card."},{"tier":"2 cards: 32 GB Blackwell (NVFP4 LLM) + ≥16 GB (ASR)","fits":true,"notes":"Smaller-box option: LLM alone on the 32 GB card with a short context (16k) and 1–2 live sessions, Voxtral (8.3 GB BF16 weights) on the second card. Estimate from the lineup page; not measured."}],"latency":[{"lane":"transcript (first text)","typical_ms":1400,"source":"measured on our server 2026-09-23 (decosa-api ops/record-all.json, field scripts: 800 and 2000 ms)"},{"lane":"checklist","typical_ms":3850,"source":"measured on our server 2026-09-23, gateway route; mean of per-script medians 3251 and 4448 ms (max 5887)"},{"lane":"issues","typical_ms":3483,"source":"measured on our server 2026-09-23, gateway route; mean of per-script medians 2518 and 4448 ms (max 5887)"},{"lane":"measurements","typical_ms":4217,"source":"measured on our server 2026-09-23, gateway route; mean of per-script medians 3985 and 4448 ms (max 5887)"},{"lane":"report (writer)","typical_ms":8500,"source":"measured on our server 2026-09-23, direct route, 8 field walk-throughs 2 at a time: median 8.5 s (7.4–11.5); gateway route, one roof replay: 9.9 s"},{"lane":"safety_sweep","typical_ms":210,"source":"measured on our server 2026-09-23, direct route, n = 8: median 0.21 s (0.16–0.24)"},{"lane":"report_check","typical_ms":2300,"source":"measured on our server 2026-09-23, direct route, n = 8: median 2.3 s (2.0–4.6), one judge call per report item, 10 in parallel; gateway route, one roof replay: 4.2 s"},{"lane":"report ready after audio ends","typical_ms":16100,"source":"measured on our server 2026-09-23, direct route, n = 8: done event a median 16.1 s (12.8–21.1) after the last transcript line; before the sweep and claim check it was 8.1 and 11.6 s on the gateway route"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"Measured latency","value":"Hosted p50 below is the time from the end of the audio to the final `done`, over 3 replays of the sample (2026-09-25, gateway route, one session at a time)."},{"label":"Typical run cost","value":"A few cents for the roof sample (a few dozen model calls) at the gateway list price. Each run shows its own measured cost."},{"label":"Data retention","value":"Audio and transcript in memory for the session only; the report comes back in done and is not stored."},{"label":"What leaves the box (self-host)","value":"Nothing, with DECOSA_LLM_ROUTE=direct."},{"label":"Outputs","value":"Structured JSON (issues with severity, quote, line and time; measurements; passed checks) and Markdown; PDF with pandoc."}],"data_handling":{"page":"/data#field","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Audio and transcript in memory for the session only; the report comes back in done and is not stored.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/operations/field","input":"mic","lanes":[{"id":"checklist","title":"Checklist","kind":"list"},{"id":"issues","title":"Issues","kind":"list"},{"id":"measurements","title":"Measurements","kind":"list"},{"id":"passed_checks","title":"Passed checks","kind":"list"},{"id":"report","title":"Inspection report","kind":"markdown"}],"samples":[{"n":1,"id":"field-roof-inspection","title":"Roof inspection","deep_link":"/tools/operations/field?sample=1&autorun=0"},{"n":2,"id":"field-electrical-panel","title":"Electrical panel check","deep_link":"/tools/operations/field?sample=2&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/field-hosted.md","selfhost":"/prompts/field-selfhost.md","assemble":"/prompts/field-assemble.md","mac":"/prompts/field-mac.md"},"rehearsal":{"bundle":"/samples/field.zip","bundle_url":"https://decosa.ai/samples/field.zip","folder":"/samples/field/","expected":"/samples/field/expected.json","files":["/samples/field/expected.json","/samples/field/inputs/roof-inspection-29s.wav","/samples/field/inputs/roof-inspection-script.json"],"bytes":680859,"checks":["the session ends with a done event","the session reports no errors","the report is titled for the site (Oak Street)","at least 2 issues are listed","the missing shingles and the lifted chimney flashing are among the issues","the roof pitch is recorded as a measurement","the overall condition is fair, poor or unsafe (not good)","at least 3 report claims are checked and supported by the transcript","every speech and model call has a signed receipt"],"licence":"Synthetic: a script written for Decosa (no real people, patients or companies) read by Decosa house voices (Kokoro-82M stock voicepacks, Apache-2.0), each allowed by the consent ledger for project decosa-field-demo. Part of decosa-api, AGPL-3.0-or-later.","about":"The first 29 seconds of a synthetic roof inspection walk-through (12 Oak Street: granule loss and three missing shingles on the south slope, chimney step flashing lifted about two inches), streamed over the live WebSocket. The session must turn it into a structured inspection report with issues and measurements, each checked against what was said.","run":{"containers":"docker compose exec api python scripts/rehearse.py field","checkout":"python scripts/rehearse.py field --bundle field.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py field"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=field","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":81.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":81.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":220,"basis":"estimate","unknown":[]},{"id":"wanted","gpu_gb":216,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":48}},"links":{"page":"/tools/operations/field","json":"/use-cases/field.json","metrics":"/metrics/field","console":"/tools/operations/field","console_sample":"/tools/operations/field?sample=1&autorun=0","stack":"/tools/operations/field#stack","try_live":"/tools/operations/field","watch":"/tools/operations/field","build":"/tools/operations/field#build","self_host":"/tools/operations/field#self-host","prompts":{"hosted":"/prompts/field-hosted.md","selfhost":"/prompts/field-selfhost.md","assemble":"/prompts/field-assemble.md","mac":"/prompts/field-mac.md"}}}