70 · Compliance and trust · live
Tariff classification memo
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Memo top-1 subheading (6-digit)145 / 200test splitn = 200held-out CBP rulings, run once; the index excludes them and their near duplicates
- Memo top-1 heading (4-digit)171 / 200test splitn = 200
- Memo top-3 subheading172 / 200test splitn = 200proposal plus up to two alternatives
- Memo top-3 heading187 / 200test splitn = 200
- Proposed (not sent to a broker)108 / 200test splitn = 200subheading right on 81% of these
- Same model without retrieval, top-1 subheading57 / 200test splitn = 200baseline: description and the GRI only
- Rulings' vote, hybrid + rerank, top-1 subheading118 / 200test splitn = 200baseline: no model, the codes of the top 5 rulings
- Rulings' vote, BM25, top-1 subheading116 / 200test splitn = 200baseline: no embedder, no reranker
- Ruling quotes verified word for word415 / 429test splitn = 429unverified quotes send the memo to a broker
Dataset
CBP CROSS New York classification rulings since 2018 (public domain), 35 headings in 11 chapters: 200 held-out test and 60 dev rulings; the index holds the other rulings minus near duplicates.
Caveats
- Queries are CBP's own product descriptions from the rulings, cut before the classification paragraphs and with codes masked; real product sheets are vaguer.
- Gold is the code CBP gave, which can be older than the 2026 HTS release; accuracy is scored at 4 and 6 digits only.
- One subset of CROSS (35 headings); rulings on the same product line by the same requester can remain in the index when their descriptions differ.
- A broker's judgment on the memos that went to a broker was not measured.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 32 s
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.004
Self-host verification
Verified on 27 Sep 2026: fresh clone into a clean directory, the api image built from it, compose up (named volume), rehearsal bundle and the heater sample against the already-running local Qwen3.8-27B (direct route) and decosa-retrieval services
9/9 rehearsal checks with the bundled 198-ruling sample set; receipts attested. The retrieval container build and the one-hour ruling fetch in the assemble prompt were not re-run in the sandbox (no new GPU loads; the fetch ran on the host).
Rehearsal bundle: tariff-classification.zip (2 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- The eval asks about products CBP already ruled on, described in CBP's own words; real product sheets are vaguer, so expect more 'needs a broker'.
- Older rulings cite statistical numbers that no longer exist; the memo keeps the valid 8- or 6-digit prefix and flags it.
- 1,992 rulings in 35 headings; a product outside those chapters gets thin evidence and goes to a broker.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- One call per memo: reads the rulings found, the HTS text of the candidate headings and the GRI, and proposes a heading, subheading and statistical number with quotes, or declines.Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Embeds the ruling set once (cached) and each product description, for the dense half of the hybrid search (BM25 is the other half).Qwen3-Embedding-0.6BApache-2.0
- Scores the 40 best chunks for each description, so the rulings the model reads are the closest products, not the closest words.Qwen3-Reranker-4BApache-2.0
- Codes against the HTS release, codes nest, quotes word for word with byte offsets, a cited ruling at the proposed heading, rulings agree, evidence not thin; the signed record.Checks and record (decosa-api, Python)AGPL-3.0-or-later
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
How often is the proposed code right?
- Memo, top-1 heading / subheading: 86% / 72% (n = 200; top-3 subheading 86%)
- Proposed memos only: 81% subheading right (108 of 200 proposed; the rest sent to a broker with the reason)
- Same model, no retrieval: 50% / 28% (heading / subheading, top-1)
- Rulings' vote, no model (hybrid + rerank / BM25): 59% / 58% (subheading, top-1)
Source: decosa-api docs/evals/tariff-classification.md, 2026-09-27; docs/evals/tariff-classification/test-score.json
Lite · smaller reranker (1)
- held-out accuracy with the 0.6B reranker: not measured yet
Standard · the hosted demo (6)
- memo top-1 subheading (6-digit), held-out rulings (n=200, run once): 145 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
- memo top-1 heading (4-digit), held-out: 171 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
- memo top-3 subheading, held-out: 172 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
- memos proposed (not sent to a broker) and their subheading accuracy, held-out: 108 / 200 proposed; 81% rightdecosa-api docs/evals/tariff-classification.md, 2026-09-27
- same model without retrieval, top-1 subheading, held-out: 57 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
- ruling vote alone (hybrid + rerank) / BM25 alone, top-1 subheading, held-out: 118 / 200 / 116 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27