Eval: disclosed virtual staging (68)
Run on our server, 26 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B through the model gateway, with
receipts), under load from other workloads. Script: scripts/staging_eval.py (steps maps, clean, plant, check,
score, label); raw results in <internal path> on our server (results.jsonl,
clean-runs.jsonl, labels.json, label-*.json).
What was measured
- Structural check: does
mode: "check"flag a staged image whose property was changed, and leave an honest staging alone? - Label read-back: does the burned-in "Virtually staged" statement and the original's link survive delivery, resizing to 1024 px and JPEG quality 70, read by Tesseract?
- Staging: how often does a first render pass the check, and what do the passing stagings look like?
Data
8 empty-room photos, all CC0:
- Poly Haven HDRI panoramas, cut to 1344x896 perspective views by us: Small Empty House (Greg Zaal; two views:
house-living,house-kitchen), Old Room (Sergej Majboroda;old-room-a,old-room-b), Small Empty Room 1 and 4 (Sergej Majboroda;room-grey,room-marble). https://polyhaven.com/a/<id> - Wikimedia Commons, SmashingIt99, "Empty apartment in Berlin with fitted kitchen and chair" 2 and 3 (1280x960;
berlin-living,berlin-open). - The beach used for "view replaced": Poly Haven Fish Hoek Beach (Greg Zaal, Rico Cilliers), CC0.
Split by room, fixed before any check ran: dev = house-kitchen, room-grey, berlin-open; test = house-living,
old-room-a, old-room-b, room-marble, berlin-living.
Clean stagings (negatives): each room staged with 4 seeds (101, 202, 303, 404), one attempt each, through the branch
server. The 18 first attempts that passed the stage-mode check were looked at by the building agent: 17 labelled clean
(furniture only), 1 not clean (old-room-b s303: a lighter rectangle of redrawn floor around a lounge chair).
house-kitchen and old-room-a produced no passing staging (see Staging below), so they have no negatives or planted edits.
Planted edits (positives), applied to each clean staging with Pillow (decosa_api/verticals/staging/plant.py):
window_removed: the largest window painted over with the wall colour beside it;wall_recoloured: pixels close to the dominant wall colour above the floor line tinted pale blue;view_replaced: the inner 80% of the largest window replaced with the beach;crack_hidden: a hairline crack drawn on plain wall in the original, the staged image left without it; in rooms whose map has damage, the largest damage box painted out of the staged image instead. Rooms whose map has no window get no window or view edit (berlin-open).
Results
Held-out test (5 rooms; nothing changed after it ran except the wording of one message, see Changes)
| cases | flagged (fail) | right category | |
|---|---|---|---|
| Clean stagings (false alarms) | 10 | 0 | |
| Window removed | 10 | 10 | 10 |
| Wall recoloured | 10 | 10 | 9 |
| View replaced | 10 | 10 | 10 |
| Crack or damage hidden | 10 | 8 | 8 |
| All planted edits | 40 | 38 | 37 |
By layer (the same test cases):
| pixels alone | model alone | |
|---|---|---|
| Planted edits flagged (of 40) | 38 | 29 |
| Clean stagings flagged (of 10) | 0 | 0 |
- The two misses are hairline cracks on the marble-effect feature wall of
room-marble: the crack is lost in the wall's own pattern. Hairline cracks overall: 5 of 7 (all 4 onberlin-living, 1 of 3 onroom-marble). - The 3 "damage hidden" cases on
old-room-bare easy: the model's damage box covered most of the wall, so painting it out changed a large area. They are counted above but say little about subtle damage. - The wall recolour with the wrong category (
berlin-livings101) was read as "other" by the pixels, with a model-only "walls" warning: the verdict was still fail. - The not-clean staging (
old-room-bs303, a redrawn floor patch) passed: the floor-pattern rule looks for a lost pattern (a border, tile joints), and this floor had little pattern. 0 of 1. - The model layer alone misses more (29 of 40, notably 5 of 10 window removals and 3 of 10 cracks): the pixel check does most of the work; the model adds categories and a second opinion.
Dev (3 rooms; thresholds were tuned here)
Clean 7 of 7 passed; window removed 4/4, wall recoloured 7/7, view replaced 4/4, crack hidden 5/7 (hairline cracks on the
grey curtained wall of room-grey).
Label read-back (Tesseract 5, label.read_back)
Every clean staging labelled with a fresh pair id and https://decosa.ai/staging/p/sp_<12>, 3 ids each:
| delivered JPEG | 1024 px wide | 1024 px, JPEG q70 | |
|---|---|---|---|
| Test (30 labels) | 30/30 | 30/30 | 30/30 |
| Dev (21 labels) | 21/21 | 21/21 | 21/21 |
| "ok" means both the statement "Virtually staged" and the exact pair id were read. With the longer branch-server address | |||
(http://127.0.0.1:8468/...) the dev numbers were 21/21, 20/21, 18/21: a long link shrinks the second line. In the check |
|||
| runs themselves (branch-server links), 13 of 13 test labels read back; on dev 6 of 8, before the pair-id alphabet (no | |||
| g/q/9, s/5, 7...) and the OCR scaling were fixed there. The label is drawn and delivered either way; the read-back is | |||
| reported, not required. |
Staging (stage mode, first attempt, 4 seeds per room)
| room | first attempt passed | why not |
|---|---|---|
| berlin-living | 4/4 | |
| old-room-b | 4/4 | |
| room-grey | 4/4 | |
| berlin-open | 3/4 | no piece clear of the doorways |
| room-marble | 3/4 | no piece clear of the fixtures |
| house-living | 0/4 | floor: the render redraws the tile border around the furniture, and the check says so |
| old-room-a | 0/4 | no piece stood clear of the windows, pipe and stains |
| house-kitchen | 0/4 | little open floor |
| all | 18/32 (56%) | a run tries up to 3 seeds, so a room that passes 3 of 4 first attempts nearly always passes |
The house-living result is the check doing its job on our own renders: the tiled border really was redrawn in all four. The floor-pattern threshold (12 blocks) was set after seeing these renders (and dev stagings), so that 4 of 4 is not a held-out number.
Honest description of the stagings: furniture is plausible and correctly placed on the floor in most passing renders
(sofas, armchairs, coffee tables, beds, side tables), with soft contact shadows. Weak points: pieces are sometimes small or
sparse (a single side table), perspective is off in a few (a sofa at a slight angle to the floor), there are faint seams
at the edge of a piece's box, and the floor right next to a piece is redrawn by the model. It is not a human stager's
quality (BoxBrownie at US$30 an image); side-by-sides are on the site (Watch) and in eval/clean/.
Cost and time (measured under load)
- Check mode: 2 image calls (map, comparison), p50 5.2 s per check on the test split; about US$0.0013 at gateway list price (2,636 prompt + ~355 generated tokens).
- Stage mode: 3 image calls (map, furniture boxes, comparison), about US$0.0021, plus one Wan2.2-VACE-Fun-A14B render (about 6 s warm, one job at a time on the shared GPU0); 15.2 s end to end for the recorded Berlin run.
- Refusals by the word patterns cost nothing (no model call).
Expected properties of the demo samples (the rehearsal bundle checks these)
check-view-replaced: statusnot_disclosed, verdictfail, avieworwindowsviolation, no pair id.check-clean: statusdisclosed, verdictpass, the label reads back (label_readback.ok), a C2PA credential (credential.kind == "c2pa"), and the pair record verifies at/record/verify(and fails onceoriginal_sha256is changed).refuse-power-linesandrefuse-bigger: HTTP 422 with categorysurroundings/size, layerpatterns, before any model call or render.- Every model call in a run has a receipt with status
signed(gateway) orattested(self-host). - The public page
/staging/p/<pair_id>names the original "ORIGINAL, UNALTERED IMAGE" and serves both files.
Caveats
- Same author for the planted edits, the checker and the clean labels; planted edits are ordinary Pillow edits, some crude (a flat-coloured window). A careful retoucher or an image model could hide edits better.
- Small: 10 clean and 40 planted cases in test, from 5 rooms (one of which has no negatives).
- Clean stagings come from our own renderer; images staged by other tools (BoxBrownie, other AI apps) were not tested.
- Some thresholds were adjusted on the dev split in three rounds (thin-line cells for cracks, the fragment rule, the hidden-element share); the floor-pattern rule was set after looking at test-room renders (see above).
- The structure map comes from the model and varies between calls; a missed window or stain weakens the rules that depend on it (the pixel regions still run).
- Not legal advice and not a compliance determination: the check is an aid, a person should look at the pair.
Changes after the test run
- One experiment (reading only floor-standing furniture as "hiding" a window) changed categories on dev and test; it was
reverted, the frozen results above are the reported ones, and the experiment's results are kept in
results-reverted-experiment.jsonl. Only the wording of that violation's message changed (painted over vs hidden).
Verdict
The check-and-disclose half is ready for a paying pilot on images staged by any tool: 38 of 40 planted edits caught with no false alarm on 10 honest stagings, the label survives portal resizing, and the pair is signed and C2PA-linked. The staging half works for rooms with clear floor and plain floors, but fails on patterned floors and cluttered rooms and is below a human stager's quality; it should be sold as a quick, honest draft, not as a BoxBrownie replacement.