Skip to content
decosa

70 · Compliance and trust · live

Tariff classification memo

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Memo top-1 subheading (6-digit)145 / 200test splitn = 200held-out CBP rulings, run once; the index excludes them and their near duplicates
  • Memo top-1 heading (4-digit)171 / 200test splitn = 200
  • Memo top-3 subheading172 / 200test splitn = 200proposal plus up to two alternatives
  • Memo top-3 heading187 / 200test splitn = 200
  • Proposed (not sent to a broker)108 / 200test splitn = 200subheading right on 81% of these
  • Same model without retrieval, top-1 subheading57 / 200test splitn = 200baseline: description and the GRI only
  • Rulings' vote, hybrid + rerank, top-1 subheading118 / 200test splitn = 200baseline: no model, the codes of the top 5 rulings
  • Rulings' vote, BM25, top-1 subheading116 / 200test splitn = 200baseline: no embedder, no reranker
  • Ruling quotes verified word for word415 / 429test splitn = 429unverified quotes send the memo to a broker

Dataset

CBP CROSS New York classification rulings since 2018 (public domain), 35 headings in 11 chapters: 200 held-out test and 60 dev rulings; the index holds the other rulings minus near duplicates.

Caveats

  • Queries are CBP's own product descriptions from the rulings, cut before the classification paragraphs and with codes masked; real product sheets are vaguer.
  • Gold is the code CBP gave, which can be older than the 2026 HTS release; accuracy is scored at 4 and 6 digits only.
  • One subset of CROSS (35 headings); rulings on the same product line by the same requester can remain in the index when their descriptions differ.
  • A broker's judgment on the memos that went to a broker was not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
27 Sep 2026
Latency, this run
n/a
p50 over passed runs
32 s
Receipts
1
Model calls
n/a
Tokens
n/a
Cost per run
$0.004

Self-host verification

Verified on 27 Sep 2026: fresh clone into a clean directory, the api image built from it, compose up (named volume), rehearsal bundle and the heater sample against the already-running local Qwen3.8-27B (direct route) and decosa-retrieval services

9/9 rehearsal checks with the bundled 198-ruling sample set; receipts attested. The retrieval container build and the one-hour ruling fetch in the assemble prompt were not re-run in the sandbox (no new GPU loads; the fetch ran on the host).

Rehearsal bundle: tariff-classification.zip (2 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
  • The eval asks about products CBP already ruled on, described in CBP's own words; real product sheets are vaguer, so expect more 'needs a broker'.
  • Older rulings cite statistical numbers that no longer exist; the memo keeps the valid 8- or 6-digit prefix and flags it.
  • 1,992 rulings in 35 headings; a product outside those chapters gets thin evidence and goes to a broker.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • One call per memo: reads the rulings found, the HTS text of the candidate headings and the GRI, and proposes a heading, subheading and statistical number with quotes, or declines.Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
  • Embeds the ruling set once (cached) and each product description, for the dense half of the hybrid search (BM25 is the other half).Qwen3-Embedding-0.6BApache-2.0
  • Scores the 40 best chunks for each description, so the rulings the model reads are the closest products, not the closest words.Qwen3-Reranker-4BApache-2.0
  • Codes against the HTS release, codes nest, quotes word for word with byte offsets, a cited ruling at the proposed heading, rulings agree, evidence not thin; the signed record.Checks and record (decosa-api, Python)AGPL-3.0-or-later

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

How often is the proposed code right?

  • Memo, top-1 heading / subheading: 86% / 72% (n = 200; top-3 subheading 86%)
  • Proposed memos only: 81% subheading right (108 of 200 proposed; the rest sent to a broker with the reason)
  • Same model, no retrieval: 50% / 28% (heading / subheading, top-1)
  • Rulings' vote, no model (hybrid + rerank / BM25): 59% / 58% (subheading, top-1)

Source: decosa-api docs/evals/tariff-classification.md, 2026-09-27; docs/evals/tariff-classification/test-score.json

Lite · smaller reranker (1)
  • held-out accuracy with the 0.6B reranker: not measured yet
Standard · the hosted demo (6)
  • memo top-1 subheading (6-digit), held-out rulings (n=200, run once): 145 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
  • memo top-1 heading (4-digit), held-out: 171 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
  • memo top-3 subheading, held-out: 172 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
  • memos proposed (not sent to a broker) and their subheading accuracy, held-out: 108 / 200 proposed; 81% rightdecosa-api docs/evals/tariff-classification.md, 2026-09-27
  • same model without retrieval, top-1 subheading, held-out: 57 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27
  • ruling vote alone (hybrid + rerank) / BM25 alone, top-1 subheading, held-out: 118 / 200 / 116 / 200decosa-api docs/evals/tariff-classification.md, 2026-09-27

How we measure · All tools