Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: disclosed virtual staging (68)

Run on our server, 26 Sep 2026, branch the pre-release branch, gateway route (Qwen3.8-27B through the model gateway, with receipts), under load from other workloads. Script: scripts/staging_eval.py (steps maps, clean, plant, check, score, label); raw results in <internal path> on our server (results.jsonl, clean-runs.jsonl, labels.json, label-*.json).

What was measured

  1. Structural check: does mode: "check" flag a staged image whose property was changed, and leave an honest staging alone?
  2. Label read-back: does the burned-in "Virtually staged" statement and the original's link survive delivery, resizing to 1024 px and JPEG quality 70, read by Tesseract?
  3. Staging: how often does a first render pass the check, and what do the passing stagings look like?

Data

8 empty-room photos, all CC0:

  • Poly Haven HDRI panoramas, cut to 1344x896 perspective views by us: Small Empty House (Greg Zaal; two views: house-living, house-kitchen), Old Room (Sergej Majboroda; old-room-a, old-room-b), Small Empty Room 1 and 4 (Sergej Majboroda; room-grey, room-marble). https://polyhaven.com/a/<id>
  • Wikimedia Commons, SmashingIt99, "Empty apartment in Berlin with fitted kitchen and chair" 2 and 3 (1280x960; berlin-living, berlin-open).
  • The beach used for "view replaced": Poly Haven Fish Hoek Beach (Greg Zaal, Rico Cilliers), CC0.

Split by room, fixed before any check ran: dev = house-kitchen, room-grey, berlin-open; test = house-living, old-room-a, old-room-b, room-marble, berlin-living.

Clean stagings (negatives): each room staged with 4 seeds (101, 202, 303, 404), one attempt each, through the branch server. The 18 first attempts that passed the stage-mode check were looked at by the building agent: 17 labelled clean (furniture only), 1 not clean (old-room-b s303: a lighter rectangle of redrawn floor around a lounge chair). house-kitchen and old-room-a produced no passing staging (see Staging below), so they have no negatives or planted edits.

Planted edits (positives), applied to each clean staging with Pillow (decosa_api/verticals/staging/plant.py):

  • window_removed: the largest window painted over with the wall colour beside it;
  • wall_recoloured: pixels close to the dominant wall colour above the floor line tinted pale blue;
  • view_replaced: the inner 80% of the largest window replaced with the beach;
  • crack_hidden: a hairline crack drawn on plain wall in the original, the staged image left without it; in rooms whose map has damage, the largest damage box painted out of the staged image instead. Rooms whose map has no window get no window or view edit (berlin-open).

Results

Held-out test (5 rooms; nothing changed after it ran except the wording of one message, see Changes)

cases flagged (fail) right category
Clean stagings (false alarms) 10 0
Window removed 10 10 10
Wall recoloured 10 10 9
View replaced 10 10 10
Crack or damage hidden 10 8 8
All planted edits 40 38 37

By layer (the same test cases):

pixels alone model alone
Planted edits flagged (of 40) 38 29
Clean stagings flagged (of 10) 0 0
  • The two misses are hairline cracks on the marble-effect feature wall of room-marble: the crack is lost in the wall's own pattern. Hairline cracks overall: 5 of 7 (all 4 on berlin-living, 1 of 3 on room-marble).
  • The 3 "damage hidden" cases on old-room-b are easy: the model's damage box covered most of the wall, so painting it out changed a large area. They are counted above but say little about subtle damage.
  • The wall recolour with the wrong category (berlin-living s101) was read as "other" by the pixels, with a model-only "walls" warning: the verdict was still fail.
  • The not-clean staging (old-room-b s303, a redrawn floor patch) passed: the floor-pattern rule looks for a lost pattern (a border, tile joints), and this floor had little pattern. 0 of 1.
  • The model layer alone misses more (29 of 40, notably 5 of 10 window removals and 3 of 10 cracks): the pixel check does most of the work; the model adds categories and a second opinion.

Dev (3 rooms; thresholds were tuned here)

Clean 7 of 7 passed; window removed 4/4, wall recoloured 7/7, view replaced 4/4, crack hidden 5/7 (hairline cracks on the grey curtained wall of room-grey).

Label read-back (Tesseract 5, label.read_back)

Every clean staging labelled with a fresh pair id and https://decosa.ai/staging/p/sp_<12>, 3 ids each:

delivered JPEG 1024 px wide 1024 px, JPEG q70
Test (30 labels) 30/30 30/30 30/30
Dev (21 labels) 21/21 21/21 21/21
"ok" means both the statement "Virtually staged" and the exact pair id were read. With the longer branch-server address
(http://127.0.0.1:8468/...) the dev numbers were 21/21, 20/21, 18/21: a long link shrinks the second line. In the check
runs themselves (branch-server links), 13 of 13 test labels read back; on dev 6 of 8, before the pair-id alphabet (no
g/q/9, s/5, 7...) and the OCR scaling were fixed there. The label is drawn and delivered either way; the read-back is
reported, not required.

Staging (stage mode, first attempt, 4 seeds per room)

room first attempt passed why not
berlin-living 4/4
old-room-b 4/4
room-grey 4/4
berlin-open 3/4 no piece clear of the doorways
room-marble 3/4 no piece clear of the fixtures
house-living 0/4 floor: the render redraws the tile border around the furniture, and the check says so
old-room-a 0/4 no piece stood clear of the windows, pipe and stains
house-kitchen 0/4 little open floor
all 18/32 (56%) a run tries up to 3 seeds, so a room that passes 3 of 4 first attempts nearly always passes

The house-living result is the check doing its job on our own renders: the tiled border really was redrawn in all four. The floor-pattern threshold (12 blocks) was set after seeing these renders (and dev stagings), so that 4 of 4 is not a held-out number.

Honest description of the stagings: furniture is plausible and correctly placed on the floor in most passing renders (sofas, armchairs, coffee tables, beds, side tables), with soft contact shadows. Weak points: pieces are sometimes small or sparse (a single side table), perspective is off in a few (a sofa at a slight angle to the floor), there are faint seams at the edge of a piece's box, and the floor right next to a piece is redrawn by the model. It is not a human stager's quality (BoxBrownie at US$30 an image); side-by-sides are on the site (Watch) and in eval/clean/.

Cost and time (measured under load)

  • Check mode: 2 image calls (map, comparison), p50 5.2 s per check on the test split; about US$0.0013 at gateway list price (2,636 prompt + ~355 generated tokens).
  • Stage mode: 3 image calls (map, furniture boxes, comparison), about US$0.0021, plus one Wan2.2-VACE-Fun-A14B render (about 6 s warm, one job at a time on the shared GPU0); 15.2 s end to end for the recorded Berlin run.
  • Refusals by the word patterns cost nothing (no model call).

Expected properties of the demo samples (the rehearsal bundle checks these)

  1. check-view-replaced: status not_disclosed, verdict fail, a view or windows violation, no pair id.
  2. check-clean: status disclosed, verdict pass, the label reads back (label_readback.ok), a C2PA credential (credential.kind == "c2pa"), and the pair record verifies at /record/verify (and fails once original_sha256 is changed).
  3. refuse-power-lines and refuse-bigger: HTTP 422 with category surroundings / size, layer patterns, before any model call or render.
  4. Every model call in a run has a receipt with status signed (gateway) or attested (self-host).
  5. The public page /staging/p/<pair_id> names the original "ORIGINAL, UNALTERED IMAGE" and serves both files.

Caveats

  • Same author for the planted edits, the checker and the clean labels; planted edits are ordinary Pillow edits, some crude (a flat-coloured window). A careful retoucher or an image model could hide edits better.
  • Small: 10 clean and 40 planted cases in test, from 5 rooms (one of which has no negatives).
  • Clean stagings come from our own renderer; images staged by other tools (BoxBrownie, other AI apps) were not tested.
  • Some thresholds were adjusted on the dev split in three rounds (thin-line cells for cracks, the fragment rule, the hidden-element share); the floor-pattern rule was set after looking at test-room renders (see above).
  • The structure map comes from the model and varies between calls; a missed window or stain weakens the rules that depend on it (the pixel regions still run).
  • Not legal advice and not a compliance determination: the check is an aid, a person should look at the pair.

Changes after the test run

  • One experiment (reading only floor-standing furniture as "hiding" a window) changed categories on dev and test; it was reverted, the frozen results above are the reported ones, and the experiment's results are kept in results-reverted-experiment.jsonl. Only the wording of that violation's message changed (painted over vs hidden).

Verdict

The check-and-disclose half is ready for a paying pilot on images staged by any tool: 38 of 40 planted edits caught with no false alarm on 10 honest stagings, the label survives portal resizing, and the pair is signed and C2PA-linked. The staging half works for rooms with clear floor and plain floors, but fails on patterned floors and cluttered rooms and is below a human stager's quality; it should be sold as a quick, honest draft, not as a BoxBrownie replacement.