Tariff classification memo (70): eval on held-out CBP rulings (27 Sep 2026)
Question: given only a product description, how often does the memo propose the code CBP gave, and how often does it know to send the product to a broker instead?
Data
- Ruling set: 1,992 CBP CROSS New York rulings (public domain, 17 U.S.C. 105), fetched 27 Sep 2026 with
scripts/tariff_fetch.py: classification rulings since 2018, found by searching 35 headings in 11 chapters (39, 42, 61, 62, 63, 64, 73, 84, 85, 94, 95) plus product words, kept only when every non-chapter-99 tariff number in the ruling shares one 6-digit subheading (one product, one answer), and not revoked or modified. Up to 70 per heading. The address block and the salutation are removed before storing; 549 rulings whose line breaks CROSS returned as runs of spaces were re-fetched and their paragraphs restored. Set hash (rulings_sha256)e9ade1b0af8f…; manifest under<internal path>. - HTS: USITC REST export of the 11 chapters and the General Rules of Interpretation, 2026 Revision 19.
- Split (
docs/evals/tariff-classification/split.json, by sha256 of the ruling number, fixed before any model run): dev 60, test 200, from the 1,861 rulings with a product description of 150+ characters. - Leakage control. The index the eval searches holds neither the held-out rulings nor any ruling whose description is a near duplicate of one (word 5-gram Jaccard 0.35 or more): 24 rulings dropped, 1,708 indexed. Rulings on similar products by the same requester stay in when their wording differs; that is also how the tool is used.
- Query: the ruling's own product description: the paragraphs after the salutation up to the first one that discusses classification (the HTSUS, a heading, the GRI, "is classified"), with any tariff number masked, capped at 2,000 characters for the search. The subject line and the holding are never in the query.
- Gold: the code CBP gave. Rulings from 2018 on may cite statistical suffixes that no longer exist; accuracy is scored at 4 and 6 digits only.
Systems
- Memo: hybrid search (Qwen3-Embedding-0.6B + BM25, reciprocal-rank fusion) over the ruling chunks, Qwen3-Reranker-4B on the top 40, up to 6 rulings with their matched chunks and holding; a second search over the HTS heading texts; one Qwen3.8-27B call (gateway route, temperature 0) with the rulings, the heading texts and the GRI; the checks. Top-3 = the proposal and up to two alternatives.
- Rulings' vote: no model: the codes of the top 5 retrieved rulings weighted 1/rank (hybrid + rerank; hybrid without rerank; BM25 only).
- No retrieval: the same model with the description only, asked for a code and two alternatives.
Results
| Held-out test, n = 200 (run once) | Heading top-1 | Heading top-3 | Subheading top-1 | Subheading top-3 |
|---|---|---|---|---|
| Memo (full pipeline) | 0.855 | 0.935 | 0.725 | 0.860 |
| Rulings' vote, hybrid + rerank | 0.780 | 0.930 | 0.590 | 0.815 |
| Rulings' vote, hybrid, no rerank | 0.790 | 0.915 | 0.585 | 0.765 |
| Rulings' vote, BM25 only | 0.780 | 0.930 | 0.580 | 0.780 |
| Same model, no retrieval | 0.500 | 0.660 | 0.285 | 0.430 |
- Proposed vs sent to a broker (test): 108 of 200 memos came back
proposed; on those the heading was right 90.7% and the subheading 80.6% (top-3 89.8%). 92 went to a broker: 70 because the model declined (a missing fact: fibre shares, gender, essential character, the use), the rest on a failed check (thin evidence 15, an unverified quote 13, no cited ruling at the proposed heading 9, a code not in the release 7; a memo can fail more than one). The model's best guess on those was right at the subheading 63% of the time: declining costs coverage, as intended. - Quotes: 415 of 429 ruling quotes were found word for word in the retrieved text (with ellipses, every fragment in order); the 14 others sent their memos to a broker.
- Retrieval: a ruling with the gold heading was among the 6 retrieved for 94% of test queries, with the gold subheading for 84.5%.
- Dev (60, used for the choices below): memo heading/subheading top-1 0.950 / 0.850; rulings' vote 0.750 / 0.583; BM25 vote 0.767 / 0.500; no retrieval 0.583 / 0.400. 34 of 60 proposed.
What was tuned on dev
- Quotes with ellipses: the first dev run showed the model quoting holdings as "The applicable subheading for ... will be ..."; the check now verifies each fragment in order instead of failing the quote.
- Old statistical numbers: a proposed 10-digit number missing from the 2026 release is cut to its longest valid prefix and flagged (non-blocking) rather than failing the memo.
- The thin-evidence threshold (0.30 on the reranker score) was set before the dev run; the dev sweep showed coverage flat from 0 to 0.65, so it was left alone. The prompt was not changed after the first dev run.
Reading the numbers
- Retrieval is most of the gain. The same model without rulings gets the subheading 28.5% of the time; with them, 72.5%. The model also beats the rulings' own vote (59%) because it reads the product against the heading texts.
- The reranker did not help the vote on test (0.590 vs 0.585 without it, 0.580 for BM25), though it did on dev. The rulings here are long and keyword-rich, so BM25 is a strong baseline; what the reranker changes is which passages the model reads, which this eval does not isolate.
- Queries are easy in one way and hard in another. They are CBP's own careful descriptions (easier than a sourcing sheet), but they are cut at the first mention of classification, which sometimes removes facts the ruling relied on.
Honest caveats
- One subset of CROSS (35 headings, NY rulings only). Products outside those headings get thin evidence.
- Gold labels are CBP's codes at the time of the ruling; some headings were renumbered since (e.g. 9405 in 2022).
- No broker reviewed the memos; "right" means "matches CBP's code", not "a broker would sign it".
- Retrieval models share GPU0 with other services; latency varied from 5 s to 35 s per memo in the runs recorded.
Expected properties of the sample runs (rehearsal/tariff-classification)
- The
wall-heatersample (a wall-mounted, hard-wired fan heater) is proposed under heading 8516, subheading 8516.29, with at least one verified quote from a retrieved ruling. - The
vague-bagsample ("a bag for carrying things") comes backneeds_broker. - Every ruling quote marked
okhas byte offsets, and the search receipt names the same index hash as the result. - Every model call has a signed (or attested) receipt, and the memo record verifies at
POST /record/verify. - The retrieved rulings for the heater include at least one CBP ruling classified in 8516.
Reproduce
python scripts/tariff_fetch.py --out <internal path> # ~1 h at 2 requests/s
python scripts/tariff_eval.py split
python scripts/tariff_eval.py run --split dev --work ~/.cache/tariff-eval/v2 # needs decosa-retrieval and the gateway
python scripts/tariff_eval.py run --split test --work ~/.cache/tariff-eval/v2
python scripts/tariff_eval.py score --split test --work ~/.cache/tariff-eval/v2
Scores: docs/evals/tariff-classification/{dev,test}-score.json. Cost: one model call per memo, about 11,000 prompt and
500 generated tokens (from the gateway receipts of the recorded runs), about $0.004 at list price.