Eval: split-sheet and metadata checker (39)
Run on our server, 25 Sep 2026, against the pre-release server (127.0.0.1:8439) on the hosted route: Qwen3.8-27B NVFP4 through
the model gateway, temperature 0, at most 2 model calls in flight. Script: scripts/splits_eval.py; data generator:
scripts/splits_data.py; metrics: docs/evals/split-sheet-check/{dev,test}-metrics.json, per-catalog issues:
{dev,test}-results.json. Cost: $0.083 of list-price tokens for the whole test set.
Data
- Synthetic, fictional catalogs (no real songs, writers or registrations): each is a 2-3 song release with a split sheet per song, a release message (DDEX ERN 4.3, validated against DDEX's published XSD, or a distributor CSV in a quarter of catalogs), a society registration export (CSV) and 0-2 co-publishing contract excerpts. Split sheets come in five styles: a filled-in form, a text table, prose, an email thread, and (about a quarter) a handwriting-style scan: the form drawn in Patrick Hand or Kalam (SIL OFL), rotated, noised and blurred, saved as an image-only PDF that the server must OCR (Tesseract 5). These are handwriting fonts, not real handwriting.
- Planted errors (0-3 per catalog, about 20% of catalogs clean): an ISWC, IPI or UPC check digit broken, an ISRC made malformed, one writer's share raised by 5 points on a split sheet (105%), a co-writer removed from the release or the registration, a composer/lyricist swap, a society changed in one source, a publisher removed from the split sheet and the registration for a writer whose contract or release names one.
- Dev: seed 39001, 8 catalogs, used while fixing the prompt (society and IPI in one column), OCR handling (misread society names and share digits, an OCR'd identifier that matches a valid one elsewhere shown as a probable misread) and the generator. Test: seed 39777, 30 catalogs, 43 planted errors, run once after dev was frozen. Nothing was changed after the test run.
1. Planted-error detection (test, 30 catalogs)
| Planted | Found |
|---|---|
| 105% split | 5 of 5 |
| Missing co-writer | 6 of 6 |
| Composer/lyricist swap | 8 of 8 |
| Society mismatch | 5 of 5 |
| Missing publisher | 1 of 1 |
| ISWC check digit | 3 of 3 |
| IPI check digits | 4 of 4 |
| UPC check digit | 5 of 5 |
| Malformed ISRC | 6 of 6 |
| All | 43 of 43 |
Precision (an issue that matches a planted error, over those plus false alarms; knock-on issues such as the 105% writer's share also differing from the registration are counted separately, 9 in all): 64.8% overall, but split by input it is two different products:
- Catalogs with typed split sheets only (15): 19 of 19 issues are planted errors; no false alarm.
- Catalogs with a scanned split sheet (15): 27 of 52. Every one of the 25 false alarms traces to OCR: identifiers
misread (
$2239359233, a digit dropped), a society read asBM), a publisher read asCoppertine Songs, a first name read asOlawaseun(so the writer looks missing). 22 of the 25 are warnings already marked "read by OCR from a scan: check the scan itself". On the 7 clean catalogs there were 3 false alarms, all on scans.
2. Extraction from the split sheets (test)
| Field | Typed (162 writers) | Handwriting-style scan via OCR (59) |
|---|---|---|
| Name | 100% | 98.3% |
| Share | 100% | 96.6% |
| Role | 93.8% | 98.3% |
| IPI | 100% | 33.9% |
| Society | 100% | 96.6% |
| Publisher | 100% | 94.9% |
The typed-role misses are all role words the table does not know ("beat + chords", "wrote the lyrics"); those go to the typed role question, which answered 10 of 10 correctly, so the comparison used the right role. Values the model reported but code could not find in the document: 0 on typed sheets. IPIs on scans are the weak spot: Tesseract confuses 5/S/$ and 1/l/! in long digit strings, and a check digit catches the damage, which is why an OCR'd identifier is a warning, not an error.
3. Arithmetic and identifiers (test)
Every share sum and every identifier in every report was recomputed by an independent implementation in the eval script (not imported from the checker): 0 disagreements on 151 share sums and 374 identifiers. The model never adds or checks anything; shares are exact fractions (33 1/3 x 3 is 100, 33.33 x 3 is a rounding note).
4. Cost and time
- Test: 143 model calls (145 receipts), 98.6k prompt and 35.5k generated tokens: $0.0028 per release at the gateway's list price ($0.30 / $1.50 per million). Median 9 s per release, max 119 s (the slow ones while other agents saturated the shared gateway).
- Demo sample (7 documents, 5 calls): 6.6-7.7 s hosted with the gateway quiet (3 smoke runs), 42-63 s under load; 4.5 s self-hosted on the direct route.
Limits of this eval
- Synthetic data from our own generator, with our own formats: real split sheets, real society exports and real DDEX feeds from distributors vary more. The generator and the checker were written by the same agent, so the structured formats are friendly to the parser; the messy text was read by the model, not by rules.
- Handwriting fonts are not handwriting. Real handwritten sheets will do worse through Tesseract.
- One run; model output at temperature 0 on a shared gateway can still vary slightly between runs.