Skip to content
decosa

97 · Any industry · Finance and insurance · preview

Fill a form from your papers

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 28 Sep 2026

  • PDF forms: wrong answers / answers written1 / 491 (0.20%)test splitn = 49150 held-out synthetic PDF forms (36 fillable, 8 flat, 6 scanned), gateway route; 95% CI 0.04-1.14%. The one: a previous street address written with its city. Scored from the output PDF, not the model.
  • PDF forms: answers the papers give, filled490 / 497 (98.6%)test splitn = 4976 left blank: a parent and child with the same name made one form ambiguous (5), and one answer cited the wrong line (blocked).
  • PDF forms: signature, signing-date and certification boxes left alone150 / 150test splitn = 150Required 100%. Checked again at the write step, whatever a caller passes.
  • PDF forms: blanks with the right reason409 / 409test splitn = 409'Your papers do not say', 'you fill this' (ID, bank), 'only you can sign or confirm'. The consent rule for Spanish and French was added after the first test run (it changed reasons, not answers).
  • Blank boxes found on flat PDFs (vector / scanned)144 / 144 and 114 / 114test splitn = 258Labels right on 144 / 144 and 113 / 114; a blank we cannot name is listed and never written. Synthetic forms from our own generator: real flat forms will be harder.
  • Wrong values / values typed6 / 1,166 (0.51%)test splitn = 1,16640 fictional forms in 5 languages x 3 source bundles (120 runs); 95% CI 0.24-1.12%. Re-run 28 Sep on the shared engine, gateway route; before the move it was 5 / 1,167. 2 prior street addresses with the city, 2 dates of birth, 1 phone, 1 comment.
  • Sourced values filled1,160 / 1,175 (98.7%)test splitn = 1,17510 left empty that a document gave, 3 wrong.
  • Fields correctly left for the person647 / 649test splitn = 649Fields no document answers, sensitive fields, signatures and declarations.
  • Forms sent by the agent0 / 120test splitn = 120And 120 of 120 forced submits (a click plus a script submit after the run) held by the network layer.
  • Planted-injection forms: leaks / sent0 / 20 and 0 / 20test splitn = 2018 flagged; the 2 script attacks (a beacon to another host, an auto-submit) were stopped at the network layer.

Dataset

PDF forms: 62 synthetic PDFs from scripts/formfill_suite_gen.py (seeded; 12 dev, 50 test): claim, benefits, W-9 style and school forms in English, Spanish and French, fillable or flat (vector or scanned). Web forms: 50 fictional forms in 5 languages x 3 synthetic source bundles plus 20 planted-injection variants (scripts/fillstop_suite_gen.py); MiniWoB++ (MIT) for the browser engine.

Caveats

  • The same author wrote the PDF generator, the forms, the detector and the checker; real forms have messier layouts. On three real federal PDFs (IRS W-9 and 8822, GSA SF-95) most labels read well after we improved the caption rules, but some still did not (not scored).
  • 4 of 5 wrong answers in the first PDF test run were our own label errors (a letter gave the address the truth said it did not); labels fixed and rescored, then the whole split re-run with the shipped code.
  • The same author wrote the generator, the forms and the checker; the forms are synthetic and short-valued (median 15 fields).
  • Uploads were not scored (no bundle has a document that is itself the invoice).
  • MiniWoB++, injection and timing runs used the direct model route (gateway out of credit); the test split used the gateway.
  • No human timed a manual fill; the typing-time comparison is a keystroke-level estimate.
  • Two blind cold-user reviews (Claude Code sub-agents playing a claimant and a security lead) read one run; no real users yet.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
28 Sep 2026
Latency, this run
n/a
p50 over passed runs
14 s
Receipts
3
Model calls
n/a
Tokens
n/a
Cost per run
$0.003

Self-host verification

Verified on 28 Sep 2026: Fresh clone of the branch, docker build of the api image (default, no browser), the api on host networking against the running local Qwen3.8-27B (direct route) and document reader; then torn down (container, volume, image, clone).

PDF mode: the Harbor Mutual sample 3 times, 8.3, 8.4 and 9.9 s, 31 of 39 boxes written each time, $0.0029 at list price, the record verified by the box's own /record/verify; the copy list for the school meal sample (site class forbidden); the site table (login.gov: identity, copy). The live web-form engine was verified self-hosted with the browser build earlier on 28 Sep (rehearsal 12 of 12).

Rehearsal bundle: fill-and-stop.zip (3 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • PDF answers come from one line of your papers (or the next); an answer that combines lines is left for you.
  • Flat and scanned PDFs: we write on the page where we found the blank; a blank we could not place or name is listed on the source sheet and left empty. Tested on synthetic flat forms; real scans vary.
  • Some PDF viewers do not show letters outside Western European sets in fillable boxes (the value is stored; check it on screen).
  • The copy list needs the website's labels (pasted, or read from a screenshot); it does not see hidden fields or later pages until you paste them.
  • Numbers come from branch runs on our server (28 Sep 2026). The planted and injection splits and MiniWoB++ were re-run through the gateway (receipted) after the move onto the shared engine; the federal-form timing and the last demo recording used our server's Qwen3.8-27B directly (same weights, receipts 'unverified').
  • A value must come from one line (or the next): an answer that combines several lines (an insurer's name, address and policy number in one box) is left for you.
  • Choices it infers (a claim type from a description, 'Yes' to a police report because a report exists) are marked 'check it'; the citation then points at supporting text, not a literal answer.
  • It does not gather facts for you: the time saved is the typing and cross-checking, not the reading. Human review time was not measured.
  • Blocking the commit also blocks a site's 'Save draft' and autosave while the agent works.
  • Hosted: demo forms and pasted HTML only; logged-in sites are the self-host command-line path (a visible browser on your machine).
  • The review translation is machine translation (the language pack); short labels can still be mistranslated.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Maps each form field to the line of your documents that answers it (12 fields per call), checks page text the patterns did not settle for instructions aimed at AI agents, and reads a widget from a screenshot when the page's code cannot set itQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Reads scanned PDFs and photos of papers (a police report photographed on a phone) into numbered lines the fields can point atDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)
  • Translates the review (field labels, reasons and your cited document lines) into the person's language; the values typed into the form are never translatedHy-MT2-7B (language pack; Qwen3.8-27B for the EU languages Hy-MT2 does not cover)Apache-2.0
  • A throwaway headless Chromium per run (no cookies, no storage, no service workers, no downloads): reads the form, types and chooses in the page, holds every request that could send the form and blocks every other hostPlaywright 1.58 with Chromium (headless)Apache-2.0 (Playwright) + BSD-3-Clause (Chromium)
  • Says whether a click the model chooses would commit something (pay, delete, send, publish, submit, security) and labels the buttons left for you; the first line, with the word list and the network hold always ondecosa-commit-detector (XLM-RoBERTa-large fine-tune, our own model)Apache-2.0 (base model MIT)
  • Field reader, the value checks (the value must be in the cited line, or the same date, time, phone number or amount written another way), the sensitive-field rules, the commit words in 32 languages, the injection patterns, the review, the signed approval and the release of the held Submitdecosa-api fill-and-stop module (decosa_api/verticals/fillstop) on the shared computer-use engine (decosa_api.cu) and the flight recorder (tool 26)AGPL-3.0-or-later (the computer-use engine it runs on, decosa_api.cu, is Apache-2.0)

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

How well does the agent use a browser on its own?

  • Shared engine, typed values only from the task text (the product's rule): 73.1% (68.7-77.1%) (307 / 420 episodes, 28 Sep 2026 re-run after fill-and-stop moved onto the engine)
  • Fill-and-stop's own observer before the move (same episodes): 76.2% (71.9-80.0%) (320 / 420; paired difference -3.1 points, 95% CI -6.9 to +0.2 (not significant), mostly date pickers)
  • Earlier baseline, same group: element table + guards / pixels only: 36.7% / 61.0% (25 Sep 2026 bench)

Source: decosa-api docs/evals/fill-and-stop.md, 28 Sep 2026

Lite · text and PDF sources, English review (1)
  • planted form suite, test split (text and PDF sources): wrong values / values typed; submit presses: 6 / 1,166 (0.51%); 0decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28, gateway route
Standard · adds scans, photos and a review in 24 languages (hosted demo) (2)
  • Harbor Mutual demo (39 fields, two PDFs and a phone photo, Spanish review): fields filled / left for the person; flag caught; submits: 33 / 6; yes; 0decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28
  • 20 planted-injection forms: leaks; submit presses; flagged: 0; 0; 18 of 20 (the other 2 were script attacks stopped at the network layer)decosa-api docs/evals/fill-and-stop.md, measured on our server 2026-09-28
Wanted: the best setup · a second model on choices and widgets (1)
  • not measured yet: not measured yetnot run

How we measure · All tools