{"about":"Hardware catalog for the self-host hardware check (src/lib/hardware.ts, /self-host/hardware, /api/hardware.json). `devices` are accelerators or machines; `models` give the memory footprint of the models the stacks use, matched to stack.json components by hf_repo. Specs come from the linked datasheets; a field listed in `unverified` could not be confirmed from a fetched page on the `checked` date and is shown with that label. Prices move a lot (GPU street prices in 2026 are far above launch prices); treat them as a rough guide.","checked":"2026-09-26","page":"https://decosa.ai/self-host/hardware","how":"Fit = each GPU component's memory (a model footprint from content/hardware.json, else stack.json vram_gb, else an estimate from parameters x bytes per weight) packed onto the GPUs; renderers that load per job count once. basis says measured, stack (stated in stack.json) or estimate.","families":{"nvidia-consumer":"NVIDIA GeForce","nvidia-workstation":"NVIDIA workstation","nvidia-datacenter":"NVIDIA datacenter","amd":"AMD","apple":"Apple Silicon (unified memory)","cpu":"CPU only","cloud":"Cloud, rented by the hour"},"devices":[{"id":"rtx-3090","name":"GeForce RTX 3090","family":"nvidia-consumer","memoryGb":24,"memoryType":"GDDR6X","unified":false,"bandwidthGbs":936,"arch":"ampere","runtime":"cuda","fp8":false,"fp4":null,"maxCount":4,"weRun":false,"price":{"usd":1499,"usdPerHour":null,"kind":"launch MSRP (2020)","note":"Sold used today; used prices not checked."},"sources":[{"label":"Wikipedia: GeForce 30 series","url":"https://en.wikipedia.org/wiki/GeForce_RTX_30_series"}],"unverified":[],"aliases":["RTX 3090","3090"],"note":"No FP8 tensor cores: FP8 checkpoints run weight-only (slower) in vLLM, NVFP4 ones do not run.","sameAs":null,"memoryOptions":[]},{"id":"rtx-4090","name":"GeForce RTX 4090","family":"nvidia-consumer","memoryGb":24,"memoryType":"GDDR6X","unified":false,"bandwidthGbs":1008,"arch":"ada","runtime":"cuda","fp8":true,"fp4":null,"maxCount":4,"weRun":false,"price":{"usd":1599,"usdPerHour":null,"kind":"launch MSRP (2022)","note":"Current street and used prices not checked."},"sources":[{"label":"Wikipedia: GeForce 40 series","url":"https://en.wikipedia.org/wiki/GeForce_RTX_40_series"}],"unverified":[],"aliases":["RTX 4090","4090"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-5090","name":"GeForce RTX 5090","family":"nvidia-consumer","memoryGb":32,"memoryType":"GDDR7","unified":false,"bandwidthGbs":1792,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":4,"weRun":false,"price":{"usd":1999,"usdPerHour":null,"kind":"launch MSRP (Jan 2025)","note":"Newegg listings on 2026-09-26 were $6,900-9,990, mostly marketplace sellers."},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/"},{"label":"Newegg listings","url":"https://www.newegg.com/p/pl?d=rtx+5090"}],"unverified":[],"aliases":["RTX 5090","5090","32 GB Blackwell"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-pro-4000","name":"RTX PRO 4000 Blackwell","family":"nvidia-workstation","memoryGb":24,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":672,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":4,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-4000/"}],"unverified":[],"aliases":["RTX PRO 4000"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-pro-4500","name":"RTX PRO 4500 Blackwell","family":"nvidia-workstation","memoryGb":32,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":896,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":4,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-4500/"}],"unverified":[],"aliases":["RTX PRO 4500"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-pro-5000-48","name":"RTX PRO 5000 Blackwell 48 GB","family":"nvidia-workstation","memoryGb":48,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":1344,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":4,"weRun":false,"price":{"usd":8599,"usdPerHour":null,"kind":"retail listing, 2026-09-26","note":"Newegg: $8,599-10,450."},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5000/"},{"label":"Newegg listings","url":"https://www.newegg.com/p/pl?d=rtx+pro+5000+blackwell"}],"unverified":[],"aliases":["RTX PRO 5000"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-pro-5000-72","name":"RTX PRO 5000 Blackwell 72 GB","family":"nvidia-workstation","memoryGb":72,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":1344,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":4,"weRun":false,"price":{"usd":9999,"usdPerHour":null,"kind":"retail listing, 2026-09-26","note":"Newegg: $9,999-17,468."},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5000/"},{"label":"Newegg listings","url":"https://www.newegg.com/p/pl?d=rtx+pro+5000+blackwell"}],"unverified":[],"aliases":[],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-pro-6000","name":"RTX PRO 6000 Blackwell 96 GB","family":"nvidia-workstation","memoryGb":96,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":1792,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":8,"weRun":true,"price":{"usd":17999,"usdPerHour":null,"kind":"retail listing, 2026-09-26","note":"Newegg: $17,999 (NVIDIA-branded), recertified from about $16,400. It launched at about $8,500 in 2025 (not confirmed from a fetched page). Max-Q (300 W) and Server Edition (1,597 GB/s) variants have the same 96 GB."},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/"},{"label":"Server Edition","url":"https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/"},{"label":"Newegg listings","url":"https://www.newegg.com/p/pl?d=rtx+pro+6000+blackwell"}],"unverified":[],"aliases":["RTX PRO 6000","96 GB card","96 GB GPU","96 GB Blackwell"],"note":"What Decosa runs: two of these in our server. Every measured GPU number on this site is from this card.","sameAs":null,"memoryOptions":[]},{"id":"rtx-6000-ada","name":"RTX 6000 Ada Generation","family":"nvidia-workstation","memoryGb":48,"memoryType":"GDDR6 ECC","unified":false,"bandwidthGbs":960,"arch":"ada","runtime":"cuda","fp8":true,"fp4":null,"maxCount":4,"weRun":false,"price":{"usd":6800,"usdPerHour":null,"kind":"launch MSRP","note":null},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/rtx-6000/"}],"unverified":["bandwidth_gbs","price"],"aliases":["RTX 6000 Ada"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"rtx-a6000","name":"RTX A6000","family":"nvidia-workstation","memoryGb":48,"memoryType":"GDDR6 ECC","unified":false,"bandwidthGbs":768,"arch":"ampere","runtime":"cuda","fp8":false,"fp4":null,"maxCount":4,"weRun":false,"price":{"usd":4650,"usdPerHour":null,"kind":"launch MSRP","note":null},"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/products/workstations/rtx-a6000/"}],"unverified":["bandwidth_gbs","price"],"aliases":["A6000"],"note":"Ampere: FP8 checkpoints run weight-only (slower) in vLLM; NVFP4 ones do not run.","sameAs":null,"memoryOptions":[]},{"id":"l4","name":"L4","family":"nvidia-datacenter","memoryGb":24,"memoryType":"GDDR6","unified":false,"bandwidthGbs":300,"arch":"ada","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/l4/"}],"unverified":[],"aliases":["L4"],"note":"72 W card; low bandwidth, so large models decode slowly.","sameAs":null,"memoryOptions":[]},{"id":"l40s","name":"L40S","family":"nvidia-datacenter","memoryGb":48,"memoryType":"GDDR6 ECC","unified":false,"bandwidthGbs":864,"arch":"ada","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/l40s/"}],"unverified":[],"aliases":["L40S"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"a100-40","name":"A100 40 GB","family":"nvidia-datacenter","memoryGb":40,"memoryType":"HBM2","unified":false,"bandwidthGbs":1555,"arch":"ampere","runtime":"cuda","fp8":false,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/a100/"}],"unverified":["bandwidth_gbs"],"aliases":["A100 40"],"note":null,"sameAs":null,"memoryOptions":[]},{"id":"a100-80","name":"A100 80 GB","family":"nvidia-datacenter","memoryGb":80,"memoryType":"HBM2e","unified":false,"bandwidthGbs":2039,"arch":"ampere","runtime":"cuda","fp8":false,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/a100/"}],"unverified":[],"aliases":["A100"],"note":"2,039 GB/s is the SXM figure; PCIe is 1,935 GB/s.","sameAs":null,"memoryOptions":[]},{"id":"h100","name":"H100 80 GB (SXM)","family":"nvidia-datacenter","memoryGb":80,"memoryType":"HBM3","unified":false,"bandwidthGbs":3350,"arch":"hopper","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/h100/"}],"unverified":[],"aliases":["H100"],"note":"The H100 NVL has 94 GB at 3,900 GB/s; the PCIe card has 80 GB at about 2,000 GB/s.","sameAs":null,"memoryOptions":[]},{"id":"h200","name":"H200 141 GB","family":"nvidia-datacenter","memoryGb":141,"memoryType":"HBM3e","unified":false,"bandwidthGbs":4800,"arch":"hopper","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA product page","url":"https://www.nvidia.com/en-us/data-center/h200/"}],"unverified":[],"aliases":["H200"],"note":"4,800 GB/s is the SXM figure.","sameAs":null,"memoryOptions":[]},{"id":"b200","name":"B200 180 GB","family":"nvidia-datacenter","memoryGb":180,"memoryType":"HBM3e","unified":false,"bandwidthGbs":8000,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA DGX B200","url":"https://www.nvidia.com/en-us/data-center/dgx-b200/"},{"label":"NVIDIA HGX","url":"https://www.nvidia.com/en-us/data-center/hgx/"}],"unverified":[],"aliases":["B200"],"note":"Sold in 8-GPU HGX/DGX systems (1,440 GB in total). Datacentre Blackwell (sm_100): builds made for the RTX PRO 6000 (sm_120) may need a different engine image.","sameAs":null,"memoryOptions":[]},{"id":"b300","name":"B300 (Blackwell Ultra)","family":"nvidia-datacenter","memoryGb":262,"memoryType":"HBM3e","unified":false,"bandwidthGbs":8000,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":8,"weRun":false,"price":null,"sources":[{"label":"NVIDIA HGX","url":"https://www.nvidia.com/en-us/data-center/hgx/"},{"label":"NVIDIA GB300 NVL72","url":"https://www.nvidia.com/en-us/data-center/gb300-nvl72/"}],"unverified":["memory_gb","bandwidth_gbs"],"aliases":["B300"],"note":"NVIDIA's pages give about 262-278 GB per GPU (2.1 TB per 8 on HGX B300); 288 GB is often quoted. The fit check uses the lower figure.","sameAs":null,"memoryOptions":[]},{"id":"gb200","name":"GB200 (per Blackwell GPU)","family":"nvidia-datacenter","memoryGb":186,"memoryType":"HBM3e","unified":false,"bandwidthGbs":8000,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":72,"weRun":false,"price":null,"sources":[{"label":"NVIDIA GB200 NVL72","url":"https://www.nvidia.com/en-us/data-center/gb200-nvl72/"}],"unverified":[],"aliases":["GB200"],"note":"Rack-scale: a GB200 superchip pairs one Grace CPU with two GPUs (372 GB); NVL72 is 72 GPUs, liquid-cooled, sold as a rack. Far beyond any single use case here.","sameAs":null,"memoryOptions":[]},{"id":"mi300x","name":"Instinct MI300X","family":"amd","memoryGb":192,"memoryType":"HBM3","unified":false,"bandwidthGbs":5300,"arch":"cdna3","runtime":"rocm","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: AMD Instinct","url":"https://en.wikipedia.org/wiki/AMD_Instinct"},{"label":"vLLM ROCm install","url":"https://docs.vllm.ai/en/latest/getting_started/installation/gpu.html"}],"unverified":[],"aliases":["MI300X"],"note":"vLLM supports MI300 (gfx942) on ROCm. NVIDIA NVFP4 checkpoints do not run on AMD: use the FP8 or BF16 build. Decosa's containers are CUDA builds and have not been run on ROCm.","sameAs":null,"memoryOptions":[]},{"id":"mi325x","name":"Instinct MI325X","family":"amd","memoryGb":256,"memoryType":"HBM3E","unified":false,"bandwidthGbs":6000,"arch":"cdna3","runtime":"rocm","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: AMD Instinct","url":"https://en.wikipedia.org/wiki/AMD_Instinct"},{"label":"ROCm vLLM Docker","url":"https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference/benchmark-docker/vllm.html"}],"unverified":["memory_type"],"aliases":["MI325X"],"note":"Same ROCm path as the MI300X; FP8 or BF16 checkpoints only.","sameAs":null,"memoryOptions":[]},{"id":"mi355x","name":"Instinct MI355X","family":"amd","memoryGb":288,"memoryType":"HBM3E","unified":false,"bandwidthGbs":8000,"arch":"cdna4","runtime":"rocm","fp8":true,"fp4":"mxfp4","maxCount":8,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: AMD Instinct","url":"https://en.wikipedia.org/wiki/AMD_Instinct"},{"label":"vLLM recipe with an MXFP4 Quark checkpoint","url":"https://docs.vllm.ai/projects/recipes/en/latest/OpenAI/GPT-OSS.html"}],"unverified":[],"aliases":["MI355X"],"note":"Native MXFP4 (AMD Quark checkpoints), not NVIDIA's NVFP4: the NVFP4 builds used here need an FP8 or MXFP4 replacement. vLLM needs ROCm 7.0 or newer for MI350-series cards.","sameAs":null,"memoryOptions":[]},{"id":"rx-7900-xtx","name":"Radeon RX 7900 XTX","family":"amd","memoryGb":24,"memoryType":"GDDR6","unified":false,"bandwidthGbs":960,"arch":"rdna3","runtime":"rocm","fp8":false,"fp4":null,"maxCount":2,"weRun":false,"price":{"usd":999,"usdPerHour":null,"kind":"launch MSRP (2022)","note":null},"sources":[{"label":"Wikipedia: Radeon RX 7000 series","url":"https://en.wikipedia.org/wiki/Radeon_RX_7000_series"},{"label":"vLLM ROCm install (gfx1100 supported)","url":"https://docs.vllm.ai/en/latest/getting_started/installation/gpu.html"}],"unverified":["fp8"],"aliases":["7900 XTX"],"note":"vLLM lists RX 7900 (gfx1100) as supported on ROCm; quantisation support is narrower than on Instinct. Untested with Decosa's stacks.","sameAs":null,"memoryOptions":[]},{"id":"m3-max","name":"Apple M3 Max","family":"apple","memoryGb":128,"memoryType":"unified LPDDR5","unified":true,"bandwidthGbs":400,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: Apple M3","url":"https://en.wikipedia.org/wiki/Apple_M3"}],"unverified":["memory_options"],"aliases":["M3 Max"],"note":"MacBook Pro (2023). 400 GB/s on the 16-core chip, 300 GB/s on the 14-core one.","sameAs":null,"memoryOptions":[36,48,64,96,128]},{"id":"m4-max","name":"Apple M4 Max","family":"apple","memoryGb":128,"memoryType":"unified","unified":true,"bandwidthGbs":546,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: Apple M4","url":"https://en.wikipedia.org/wiki/Apple_M4"}],"unverified":[],"aliases":["M4 Max"],"note":"MacBook Pro (2024) and Mac Studio (2025). 546 GB/s with the 40-core GPU, 410 GB/s with the 32-core one. Apple made no M4 Ultra.","sameAs":null,"memoryOptions":[36,48,64,128]},{"id":"m3-ultra","name":"Apple M3 Ultra (Mac Studio)","family":"apple","memoryGb":512,"memoryType":"unified","unified":true,"bandwidthGbs":819,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":true,"price":null,"sources":[{"label":"Wikipedia: Mac Studio","url":"https://en.wikipedia.org/wiki/Mac_Studio"},{"label":"Decosa Mac measurements","url":"/self-host/mac"}],"unverified":[],"aliases":["M3 Ultra","Mac Studio"],"note":"The Mac every measured Mac number here comes from (512 GB). Replaced by the M5 Ultra Mac Studio in September 2026.","sameAs":null,"memoryOptions":[96,256,512]},{"id":"m5","name":"Apple M5","family":"apple","memoryGb":32,"memoryType":"unified","unified":true,"bandwidthGbs":153,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: Apple M5","url":"https://en.wikipedia.org/wiki/Apple_M5"}],"unverified":["memory_options"],"aliases":[],"note":"MacBook Pro 14, MacBook Air, iPad Pro (from October 2025).","sameAs":null,"memoryOptions":[16,24,32]},{"id":"m5-pro","name":"Apple M5 Pro","family":"apple","memoryGb":64,"memoryType":"unified","unified":true,"bandwidthGbs":307,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":null,"sources":[{"label":"Wikipedia: Apple M5","url":"https://en.wikipedia.org/wiki/Apple_M5"},{"label":"Apple Newsroom, Sep 2026","url":"https://www.apple.com/newsroom/2026/09/the-new-mac-mini-and-mac-studio-are-available-today/"}],"unverified":[],"aliases":[],"note":"MacBook Pro (March 2026) and Mac mini (September 2026).","sameAs":null,"memoryOptions":[24,48,64]},{"id":"m5-max","name":"Apple M5 Max","family":"apple","memoryGb":128,"memoryType":"unified","unified":true,"bandwidthGbs":460,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":{"usd":2499,"usdPerHour":null,"kind":"Mac Studio starting price","note":"36 GB configuration; the 128 GB price was not checked."},"sources":[{"label":"Apple: Mac Studio specs","url":"https://www.apple.com/mac-studio/specs/"},{"label":"Apple: MacBook Pro specs","url":"https://www.apple.com/macbook-pro/specs/"}],"unverified":["price"],"aliases":["M5 Max"],"note":"Mac Studio (September 2026) and MacBook Pro. 460 GB/s in the Mac Studio (32-core GPU); the 40-core GPU has 614 GB/s.","sameAs":null,"memoryOptions":[36,48,64,128]},{"id":"m5-ultra","name":"Apple M5 Ultra (Mac Studio)","family":"apple","memoryGb":512,"memoryType":"unified","unified":true,"bandwidthGbs":1200,"arch":"apple","runtime":"mlx","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":{"usd":5499,"usdPerHour":null,"kind":"Mac Studio starting price","note":"96 GB configuration; the 512 GB price was not checked."},"sources":[{"label":"Apple: Mac Studio specs","url":"https://www.apple.com/mac-studio/specs/"},{"label":"Apple Newsroom, Sep 2026","url":"https://www.apple.com/newsroom/2026/09/the-new-mac-mini-and-mac-studio-are-available-today/"}],"unverified":["price"],"aliases":["M5 Ultra"],"note":"Mac Studio from September 2026. Not measured with Decosa's stacks yet.","sameAs":null,"memoryOptions":[96,256,512]},{"id":"cpu","name":"CPU only (no GPU)","family":"cpu","memoryGb":512,"memoryType":"system RAM","unified":false,"bandwidthGbs":null,"arch":"cpu","runtime":"cpu","fp8":false,"fp4":null,"maxCount":1,"weRun":false,"price":null,"sources":[],"unverified":[],"aliases":["CPU only","Any CPU","any Linux","8-core CPU"],"note":"Rule engines, signing, OCR, embeddings and small voices run on CPU. Language models do not run at a usable speed here: point them at a remote endpoint instead.","sameAs":null,"memoryOptions":[16,32,64,128,256,512]},{"id":"cloud-h100","name":"Cloud: H100 80 GB","family":"cloud","memoryGb":80,"memoryType":"HBM3","unified":false,"bandwidthGbs":3350,"arch":"hopper","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":2.69,"kind":"on demand, per GPU-hour, 2026-09-26","note":"RunPod $2.69-3.49 (SXM), Lambda $3.99-4.29."},"sources":[{"label":"RunPod pricing","url":"https://www.runpod.io/pricing"},{"label":"Lambda pricing","url":"https://lambda.ai/pricing"}],"unverified":[],"aliases":["H100"],"note":"The H100 NVL has 94 GB at 3,900 GB/s; the PCIe card has 80 GB at about 2,000 GB/s.","sameAs":"h100","memoryOptions":[]},{"id":"cloud-h200","name":"Cloud: H200 141 GB","family":"cloud","memoryGb":141,"memoryType":"HBM3e","unified":false,"bandwidthGbs":4800,"arch":"hopper","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":3.59,"kind":"on demand, per GPU-hour, 2026-09-26","note":"RunPod $3.59-4.59."},"sources":[{"label":"RunPod pricing","url":"https://www.runpod.io/pricing"}],"unverified":[],"aliases":["H200"],"note":"4,800 GB/s is the SXM figure.","sameAs":"h200","memoryOptions":[]},{"id":"cloud-l40s","name":"Cloud: L40S 48 GB","family":"cloud","memoryGb":48,"memoryType":"GDDR6 ECC","unified":false,"bandwidthGbs":864,"arch":"ada","runtime":"cuda","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":0.79,"kind":"on demand, per GPU-hour, 2026-09-26","note":"RunPod $0.79-1.09."},"sources":[{"label":"RunPod pricing","url":"https://www.runpod.io/pricing"}],"unverified":[],"aliases":["L40S"],"note":null,"sameAs":"l40s","memoryOptions":[]},{"id":"cloud-rtx-pro-6000","name":"Cloud: RTX PRO 6000 96 GB","family":"cloud","memoryGb":96,"memoryType":"GDDR7 ECC","unified":false,"bandwidthGbs":1792,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":1.69,"kind":"on demand, per GPU-hour, 2026-09-26","note":"RunPod $1.69-2.09. The same card Decosa measures on."},"sources":[{"label":"RunPod pricing","url":"https://www.runpod.io/pricing"}],"unverified":[],"aliases":["RTX PRO 6000","96 GB card","96 GB GPU","96 GB Blackwell"],"note":"What Decosa runs: two of these in our server. Every measured GPU number on this site is from this card.","sameAs":"rtx-pro-6000","memoryOptions":[]},{"id":"cloud-b200","name":"Cloud: B200 180 GB","family":"cloud","memoryGb":180,"memoryType":"HBM3e","unified":false,"bandwidthGbs":8000,"arch":"blackwell","runtime":"cuda","fp8":true,"fp4":"nvfp4","maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":5.98,"kind":"on demand, per GPU-hour, 2026-09-26","note":"RunPod $5.98-6.79, Lambda $6.69-6.99."},"sources":[{"label":"RunPod pricing","url":"https://www.runpod.io/pricing"},{"label":"Lambda pricing","url":"https://lambda.ai/pricing"}],"unverified":[],"aliases":["B200"],"note":"Sold in 8-GPU HGX/DGX systems (1,440 GB in total). Datacentre Blackwell (sm_100): builds made for the RTX PRO 6000 (sm_120) may need a different engine image.","sameAs":"b200","memoryOptions":[]},{"id":"cloud-mi300x","name":"Cloud: MI300X 192 GB","family":"cloud","memoryGb":192,"memoryType":"HBM3","unified":false,"bandwidthGbs":5300,"arch":"cdna3","runtime":"rocm","fp8":true,"fp4":null,"maxCount":8,"weRun":false,"price":{"usd":null,"usdPerHour":2.99,"kind":"on demand, per GPU-hour, 2026-09-26","note":"Hot Aisle $2.99 (VM), $3.39 (8x bare metal)."},"sources":[{"label":"Hot Aisle pricing","url":"https://hotaisle.xyz/pricing/"}],"unverified":[],"aliases":["MI300X"],"note":"vLLM supports MI300 (gfx942) on ROCm. NVIDIA NVFP4 checkpoints do not run on AMD: use the FP8 or BF16 build. Decosa's containers are CUDA builds and have not been run on ROCm.","sameAs":"mi300x","memoryOptions":[]}],"presets":[{"id":"cpu-64x1","label":"CPU only, 64 GB RAM"},{"id":"rtx-4090x1","label":"GeForce RTX 4090"},{"id":"rtx-5090x1","label":"GeForce RTX 5090"},{"id":"rtx-5090x2","label":"2x GeForce RTX 5090"},{"id":"l40sx1","label":"L40S"},{"id":"h100x1","label":"H100 80 GB (SXM)"},{"id":"rtx-pro-6000x1","label":"RTX PRO 6000 Blackwell 96 GB"},{"id":"rtx-pro-6000x2","label":"2x RTX PRO 6000 Blackwell 96 GB"},{"id":"m3-ultra-96x1","label":"Apple M3 Ultra (Mac Studio), 96 GB"},{"id":"m5-max-64x1","label":"Apple M5 Max, 64 GB"}],"use_cases":[{"id":"clinical","name":"Visit copilot","page":"https://decosa.ai/clinics/visit-copilot#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card, live pass only","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"asr-live","role":"Pass 1: live streaming transcript for the in-visit view (no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-lite","role":"Language model for the lite tier: live lanes and the final note from the live transcript","name":"Qwen3.8-27B (official FP8)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."}]},{"id":"standard","label":"Standard · one 96 GB Blackwell card, two passes","gpu_gb":85.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Pass 1: live streaming transcript for the in-visit view (no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mapping","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr-pass2","role":"Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note cites","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"detail-checker","role":"Note self-check, detail step (M17): each drug, dose, frequency, date, side and number in a note sentence read against its transcript lines; a 'detail not in the visit' flag becomes a changed-detail error","name":"decosa-note-detail-checker-modernbert-large (M17, our own model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · adds DeepSeek V4 Flash as the note writer on 2x 96 GB","gpu_gb":277.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Pass 1: live streaming transcript for the in-visit view (no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Language model: live SOAP draft, guidance report, practitioner lanes, window role map, cited note, the self-check's sentence judge, assessment codes and the paperwork field mapping","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr-pass2","role":"Speaker labels during the visit: rolling windows (every 15 s of new audio, 6 s overlap) re-transcribed with speaker labels; the committed turns become the visit transcript the note cites","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"writer-alt","role":"Note writer alternative: the loop's best writer","name":"DeepSeek V4 Flash","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":166,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash (official FP8 + FP4 experts): About 83 GB of weights per GPU across two 96 GB cards at 32k context (clinical stack.json); little room left. The minimum is an estimate. (stack.json lists 166 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 40 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 48 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"sales","name":"Sales & meeting copilot","page":"https://decosa.ai/tools","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm_fp8","role":"Lane model, lite tier (48 GB card)","name":"Qwen3.8-27B FP8 (official)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."}]},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Lane model (objection, next question, CRM fields, follow-up email)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"best","label":"Best · two 96 GB cards for the LLM","gpu_gb":216,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm_best","role":"Lane model, best tier (2× 96 GB)","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate. (stack.json lists 178.6 GB for this component.)"}]},{"id":"wanted","label":"Wanted · the largest open flash models","gpu_gb":1344,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm_net_glm","role":"Lane model, network tier (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"llm_net_v41","role":"Lane model, network tier (wanted)","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 36 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 44 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"code","name":"Private code assistant","page":"https://decosa.ai/tools/developer/code#self-host","tiers":[{"id":"lite","label":"Lite · runs on one 32 GB Blackwell card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Coding model (chat, edit, agent tool calls)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Coding model (chat, edit, agent tool calls)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"best","label":"Best · two 96 GB cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-large","role":"Optional larger tier (two GPUs)","name":"DeepSeek-V4-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · the two most-used open coding models","gpu_gb":1320,"basis":"estimate","unknown":[],"components":[{"id":"llm-wanted","role":"Network-hosted flash model (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"dsv41-wanted","role":"Coding model, datacentre class","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"studio","name":"Decosa Studio","page":"https://decosa.ai/studio#self-host","tiers":[{"id":"lite","label":"Lite · drafts on one 24–32 GB card","gpu_gb":69.6,"basis":"estimate","unknown":["image-lite"],"components":[{"id":"music","role":"Music generation, fast drafts (text and lyrics to song)","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"image-lite","role":"Image generation, fast drafts","name":"Z-Image-Turbo","where":"gpu","load":"job","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Memory not stated in stack.json and not derivable (no parameter count)."},{"id":"video-lite","role":"Video generation, drafts","name":"Wan2.1-T2V-1.3B","where":"gpu","load":"job","gb":7.2,"min_gb":7.2,"weights_gb":5.2,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.3B parameters at 4 bytes (FP32) per weight is about 5.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"voice","role":"Text to speech (preset voices)","name":"Kokoro-82M","where":"gpu","load":"resident","gb":1.4,"min_gb":1.4,"weights_gb":0.3,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 0.082B parameters at 4 bytes (FP32) per weight is about 0.3 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"video","role":"Video generation (text to video, no audio)","name":"Wan2.1-T2V-14B","where":"gpu","load":"job","gb":68.2,"min_gb":68.2,"weights_gb":56,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 14B parameters at 4 bytes (FP32) per weight is about 56 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]},{"id":"standard","label":"Standard · the hosted demo, about 50 GB of one 96 GB card","gpu_gb":50,"basis":"stack","unknown":[],"components":[{"id":"music3","role":"Music generation (songs with vocals and lyrics)","name":"MiniMax-Music3","where":"gpu","load":"job","gb":50,"min_gb":50,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-Music3: About 50 GB free on the card for Music3 (music-gen-cleared stack.json); the MLX port peaked at 50 GB on the Mac (mac.json)."},{"id":"music","role":"Music generation, fast drafts (text and lyrics to song)","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"image","role":"Image generation","name":"Qwen-Image-2512","where":"gpu","load":"job","gb":41.6,"min_gb":41.6,"weights_gb":null,"basis":"measured","precision":null,"gpus":1,"source":"Qwen-Image-2512: Peak 41.6 GB at 1664x928, measured in the studio on 2026-09-23 (stack.json)."},{"id":"video-h3-fal","role":"Hosted video (default): 5 s clips with audio","name":"MiniMax H3 Max (via fal)","where":"remote","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"A hosted API: nothing loads on your machine."}]},{"id":"best","label":"Best · self-host only, a whole 96 GB card","gpu_gb":96,"basis":"stack","unknown":[],"components":[{"id":"music3","role":"Music generation (songs with vocals and lyrics)","name":"MiniMax-Music3","where":"gpu","load":"job","gb":50,"min_gb":50,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-Music3: About 50 GB free on the card for Music3 (music-gen-cleared stack.json); the MLX port peaked at 50 GB on the Mac (mac.json)."},{"id":"image","role":"Image generation","name":"Qwen-Image-2512","where":"gpu","load":"job","gb":41.6,"min_gb":41.6,"weights_gb":null,"basis":"measured","precision":null,"gpus":1,"source":"Qwen-Image-2512: Peak 41.6 GB at 1664x928, measured in the studio on 2026-09-23 (stack.json)."},{"id":"video-ltx","role":"Video with synchronized audio, self-host only","name":"LTX-2.5 22B distilled","where":"gpu","load":"job","gb":50,"min_gb":44,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"LTX-2 22B video: LTX-2.5 distilled renders within about 50 GB of free VRAM (measured 2026-09-23, studio stack.json); LTX-2.3 with a character LoRA about 44 GB (owner's pipeline notes, characters stack.json)."},{"id":"video-h3","role":"Video with synchronized audio and reference control, self-host only where licensed","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."}]},{"id":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","gpu_gb":96,"basis":"stack","unknown":[],"components":[{"id":"video-h3","role":"Video with synchronized audio and reference control, self-host only where licensed","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."}]}],"mac":{"fit":"partial","memory_gb":64,"tier":"standard"},"undetermined":[{"id":"image-lite","name":"Z-Image-Turbo","tiers":["lite"]}],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"MiniMax-Music3 needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 50 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 50 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"no","recommended":null,"reason":"MiniMax-Music3 needs about 50 GB on one GPU; each GPU here has 32 GB."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 50 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (106.2 of 80 GB)."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (106.2 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (106.2 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Hosted-default video (MiniMax H3 Max) is a call to fal; LTX-2.5 self-host video runs on the Mac, about 3x slower than a Blackwell card"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Hosted-default video (MiniMax H3 Max) is a call to fal; LTX-2.5 self-host video runs on the Mac, about 3x slower than a Blackwell card"}}},{"id":"field","name":"Field reports","page":"https://decosa.ai/tools/operations/field#self-host","tiers":[{"id":"lite","label":"Lite · 4-bit on a 32 GB card, plus a small card for speech","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Lanes and report (language model)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"standard","label":"Standard · one 96 GB card (the hosted demo)","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Lanes and report (language model)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"best","label":"Best · two 96 GB cards","gpu_gb":220,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr_final","role":"Final transcript pass (speaker-attributed)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm_best","role":"Lanes and report (larger model)","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · the largest open flash models","gpu_gb":216,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm_wanted","role":"Lanes and report (network-hosted flash model)","name":"GLM-5.3-Flash or DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 36 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 44 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"translate","name":"Live translation","page":"https://decosa.ai/tools/operations/translate#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":64,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-lite","role":"Translation lane model, lite tier","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Translation lane model (translation, glossary, summary)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"best","label":"Best · two 96 GB cards","gpu_gb":216,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-best","role":"Translation lane model, best tier","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · the largest open flash models","gpu_gb":1344,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition (streaming)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-network-glm","role":"Translation lane model, network-hosted (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"llm-network-dsv41","role":"Translation lane model, network-hosted (wanted)","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."}]},{"id":"alternate-specialist-mt","label":"Specialist translation models","gpu_gb":9.4,"basis":"estimate","unknown":[],"components":[{"id":"mt-hy","role":"Specialist translation model for the translation lane","name":"Hy-MT2-7B / Hy-MT2-30B-A3B-FP8","where":"gpu","load":"resident","gb":9.4,"min_gb":9.4,"weights_gb":7,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 7B parameters at 1 byte per weight is about 7 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 36 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 44 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 49.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"record","name":"Tamper-evident record","page":"https://decosa.ai/tools/developer/record#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card, captions only","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-lite","role":"Lite tier: actions lane and a cited summary written from the live captions","name":"Qwen3.8-27B (official FP8)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."}]},{"id":"standard","label":"Standard · the hosted demo, two passes","gpu_gb":85.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After the session: speaker-attributed transcript, one line per turn","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm","role":"Actions lane, speaker roles, cited summary, claim verifier","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash writes and checks","gpu_gb":220,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After the session: speaker-attributed transcript, one line per turn","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm-best","role":"Best tier: summary writer and verifier","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 40 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 48 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"provenance","name":"Provenance and consent","page":"https://decosa.ai/developers#block-content-credentials","tiers":[{"id":"lite","label":"Lite · credential and receipt, CPU only","gpu_gb":14.6,"basis":"stack","unknown":[],"components":[{"id":"c2pa","role":"Content credential signer and reader (C2PA manifests)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"music","role":"Music generation, fast drafts (stamped)","name":"ACE-Step 1.5 turbo","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."}]},{"id":"standard","label":"Standard · credential, receipt and image watermark (hosted)","gpu_gb":68.2,"basis":"estimate","unknown":[],"components":[{"id":"c2pa","role":"Content credential signer and reader (C2PA manifests)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Invisible image watermark (payload: the receipt id)","name":"TrustMark variant Q","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"image","role":"Image generation (stamped)","name":"Qwen-Image-2512","where":"gpu","load":"job","gb":41.6,"min_gb":41.6,"weights_gb":null,"basis":"measured","precision":null,"gpus":1,"source":"Qwen-Image-2512: Peak 41.6 GB at 1664x928, measured in the studio on 2026-09-23 (stack.json)."},{"id":"video","role":"Video generation (stamped)","name":"Wan2.1-T2V-14B","where":"gpu","load":"job","gb":68.2,"min_gb":68.2,"weights_gb":56,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 14B parameters at 4 bytes (FP32) per weight is about 56 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"music3","role":"Music generation (stamped)","name":"MiniMax-Music3","where":"gpu","load":"job","gb":50,"min_gb":50,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-Music3: About 50 GB free on the card for Music3 (music-gen-cleared stack.json); the MLX port peaked at 50 GB on the Mac (mac.json)."}]}],"mac":{"fit":"partial","memory_gb":64,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen-Image-2512 needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 68.2 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 68.2 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 68.2 GB of GPU memory at the smallest settings; 64 GB available. The lite tier fits."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 68.2 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (159.8 of 80 GB)."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (159.8 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (159.8 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Video renders need a CUDA GPU"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Video renders need a CUDA GPU"}}},{"id":"auditor","name":"Endpoint auditor","page":"https://decosa.ai/tools/developer/auditor#self-host","tiers":[{"id":"lite","label":"Lite · audits only, no GPU","gpu_gb":13,"basis":"stack","unknown":[],"components":[{"id":"auditor","role":"Probe runner, scorer and signer (no model; runs on CPU)","name":"decosa-api auditor (decosa_api/verticals/auditor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ref-4b","role":"Small reference model (swap and quantisation demos)","name":"Qwen3.5-4B-Base (BF16)","where":"gpu","load":"resident","gb":13,"min_gb":13,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 13 in stack.json."}]},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","gpu_gb":70.6,"basis":"stack","unknown":[],"components":[{"id":"auditor","role":"Probe runner, scorer and signer (no model; runs on CPU)","name":"decosa-api auditor (decosa_api/verticals/auditor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ref-27b","role":"Reference model (golden outputs, logprobs, noise band)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"ref-4b","role":"Small reference model (swap and quantisation demos)","name":"Qwen3.5-4B-Base (BF16)","where":"gpu","load":"resident","gb":13,"min_gb":13,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 13 in stack.json."}]},{"id":"best","label":"Best · two 96 GB cards","gpu_gb":249.6,"basis":"stack","unknown":[],"components":[{"id":"auditor","role":"Probe runner, scorer and signer (no model; runs on CPU)","name":"decosa-api auditor (decosa_api/verticals/auditor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ref-27b","role":"Reference model (golden outputs, logprobs, noise band)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"ref-dsv4","role":"Larger reference (planned)","name":"DeepSeek-V4-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · references for the most-used open models","gpu_gb":1320,"basis":"estimate","unknown":[],"components":[{"id":"auditor","role":"Probe runner, scorer and signer (no model; runs on CPU)","name":"decosa-api auditor (decosa_api/verticals/auditor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"glm-wanted","role":"Reference for GLM-5.3-Flash endpoints","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"dsv41-wanted","role":"Reference for DeepSeek-V4.1-Flash endpoints","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."}]}],"mac":{"fit":"full","memory_gb":16,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 33 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 41 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (70.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (70.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (16 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (16 of 64 GB)."}}},{"id":"ugc","name":"Disclosed UGC ads","page":"https://decosa.ai/studio/brand","tiers":[{"id":"lite","label":"Lite · fully open (Apache-2.0), on your own GPU","gpu_gb":51.4,"basis":"estimate","unknown":[],"components":[{"id":"llm_fp8","role":"Lane model, lite tier","name":"Qwen3.8-27B FP8 (official)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."},{"id":"video","role":"Lite-tier shot renderer (silent 5 s clips, 480x832), fully open weights","name":"Wan2.1-T2V-14B","where":"gpu","load":"job","gb":17.8,"min_gb":17.8,"weights_gb":14,"basis":"estimate","precision":"fp8","gpus":1,"source":"Estimate: 14B parameters at 1 byte per weight is about 14 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"voice","role":"Voice-over (preset synthetic voices only)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Content credential and render receipt (provenance kit)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Invisible watermark on every frame (payload: the receipt id)","name":"TrustMark variant Q","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted default, MiniMax H3 via fal at cost","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Lane model (brief check, hooks, script, claims review, disclosure text, storyboard)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"h3","role":"Hosted default video: presenter still, lip-synced presenter shots and reference-consistent B-roll (with audio)","name":"MiniMax H3 Max (via fal)","where":"remote","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"A hosted API: nothing loads on your machine."},{"id":"voice","role":"Voice-over (preset synthetic voices only)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Content credential and render receipt (provenance kit)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Invisible watermark on every frame (payload: the receipt id)","name":"TrustMark variant Q","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best for self-host · MiniMax H3 (licence pending) or LTX-2.3 on your own GPU","gpu_gb":207.4,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Lane model (brief check, hooks, script, claims review, disclosure text, storyboard)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"h3-selfhost","role":"Self-host: presenter stills, lip-synced shots and reference-consistent B-roll on your own GPU","name":"MiniMax-H3 (licence pending, self-host)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."},{"id":"ltx","role":"Self-host alternative: talking-head shots with audio and character LoRAs","name":"LTX-2.3 22B","where":"gpu","load":"resident","gb":53.8,"min_gb":53.8,"weights_gb":44,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 22B parameters at 2 bytes (BF16) per weight is about 44 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"voice","role":"Voice-over (preset synthetic voices only)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Content credential and render receipt (provenance kit)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Invisible watermark on every frame (payload: the receipt id)","name":"TrustMark variant Q","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","gpu_gb":153.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Lane model (brief check, hooks, script, claims review, disclosure text, storyboard)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"h3-selfhost","role":"Self-host: presenter stills, lip-synced shots and reference-consistent B-roll on your own GPU","name":"MiniMax-H3 (licence pending, self-host)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."},{"id":"voice","role":"Voice-over (preset synthetic voices only)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Content credential and render receipt (provenance kit)","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Invisible watermark on every frame (payload: the receipt id)","name":"TrustMark variant Q","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":{"fit":"partial","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Hosted-default video renders on fal (a hosted API); LTX-2.5 self-host video with audio runs on the Mac (measured in the studio, 184 s per 5 s clip), but the LTX-2.3 lip-sync path was not run on the Mac"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Hosted-default video renders on fal (a hosted API); LTX-2.5 self-host video with audio runs on the Mac (measured in the studio, 184 s per 5 s clip), but the LTX-2.3 lip-sync path was not run on the Mac"}}},{"id":"characters","name":"Animated characters","page":"https://decosa.ai/studio/characters","tiers":[{"id":"lite","label":"Lite · greetings on CPU from clips already rendered","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"voice","role":"Stock voice for the greeting (no cloning)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · new clips on about 45 GB of one 96 GB card","gpu_gb":99.2,"basis":"stack","unknown":[],"components":[{"id":"reference","role":"Character design: one fixed reference image per character (made once, reused for every clip)","name":"Qwen-Image-2512","where":"gpu","load":"job","gb":41.6,"min_gb":41.6,"weights_gb":null,"basis":"measured","precision":null,"gpus":1,"source":"Qwen-Image-2512: Peak 41.6 GB at 1664x928, measured in the studio on 2026-09-23 (stack.json)."},{"id":"video","role":"Character clips: image-to-video from the reference image (first frame + VACE reference), one clip per motion","name":"Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."},{"id":"voice","role":"Stock voice for the greeting (no cloning)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Line check: judges the greeting line with the name masked as {name}","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · self-host lip-sync with LTX-2.3 or MiniMax H3 (licence pending)","gpu_gb":96,"basis":"stack","unknown":[],"components":[{"id":"video-ltx","role":"Self-host only: talking characters with joint audio and lip-sync","name":"LTX-2.3 22B + character LoRA","where":"gpu","load":"job","gb":50,"min_gb":44,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"LTX-2 22B video: LTX-2.5 distilled renders within about 50 GB of free VRAM (measured 2026-09-23, studio stack.json); LTX-2.3 with a character LoRA about 44 GB (owner's pipeline notes, characters stack.json). (stack.json lists 44 GB for this component.)"},{"id":"voice","role":"Stock voice for the greeting (no cloning)","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"video-h3","role":"Talking-character video with native audio (reference-to-video and lip-sync modes)","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."}]}],"mac":{"fit":"partial","memory_gb":64,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen-Image-2512 needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 61.6 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 69.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 69.6 GB of GPU memory at the smallest settings; 64 GB available. The lite tier fits."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 75.2 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (133.2 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Character video (Wan 2.2) needs a CUDA GPU"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Character video (Wan 2.2) needs a CUDA GPU"}}},{"id":"grounding","name":"Grounding check","page":"https://decosa.ai/tools/developer/grounding#self-host","tiers":[{"id":"lite","label":"Lite · CPU only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: segmentation, evidence selection, gate, signed report (no model; CPU)","name":"decosa-api grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"nli","role":"Lite judge: NLI cross-encoder on CPU (self-host only)","name":"cross-encoder/nli-deberta-v3-base","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the judge (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: segmentation, evidence selection, gate, signed report (no model; CPU)","name":"decosa-api grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per sentence, verdict and cited spans","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · a panel of the largest open judges","gpu_gb":1377.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: segmentation, evidence selection, gate, signed report (no model; CPU)","name":"decosa-api grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per sentence, verdict and cited spans","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"dsv41-wanted","role":"Second judge: is each sentence supported by its sources?","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."},{"id":"glm-wanted","role":"Third judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-fast-judge","label":"Fast option-token judge","gpu_gb":57,"basis":"estimate","unknown":[],"components":[{"id":"gemma-judge","role":"Fast judge with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate. (stack.json lists 49 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"deposition","name":"Deposition and hearing digest","page":"https://decosa.ai/legal/deposition#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: digest, check and pairs on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Digest writer, cite checker and contradiction judge","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr-rough","role":"Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: digest writer and checker for many-witness matters","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr-rough","role":"Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"wanted","label":"Wanted · GLM-5.3-Flash on your own hardware","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-network-glm","role":"Multi-witness matters, on your own bigger box (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"asr-rough","role":"Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"hiring-screen","name":"Receipted hiring screen","page":"https://decosa.ai/family#autotalent","tiers":[{"id":"lite","label":"Lite · one 32 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"screen","role":"Redaction, quote check, scoring rule, decision log and impact-ratio tables (no model; CPU)","name":"decosa-api hiring screen (decosa_api/verticals/hiring)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm_home","role":"Judge on a single 32 GB card","name":"Qwen3.8-27B (NVIDIA NVFP4), short context","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"standard","label":"Standard · the hosted route","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"screen","role":"Redaction, quote check, scoring rule, decision log and impact-ratio tables (no model; CPU)","name":"decosa-api hiring screen (decosa_api/verticals/hiring)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Judge: one verdict and quote per requirement","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"best","label":"Best · two judges on one 96 GB card","gpu_gb":95.6,"basis":"estimate","unknown":[],"components":[{"id":"screen","role":"Redaction, quote check, scoring rule, decision log and impact-ratio tables (no model; CPU)","name":"decosa-api hiring screen (decosa_api/verticals/hiring)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Judge: one verdict and quote per requirement","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"judge2","role":"Second judge: requirements where the two disagree go to a person","name":"Gemma-4-31B-it","where":"gpu","load":"resident","gb":38,"min_gb":34,"weights_gb":31,"basis":"estimate","precision":"fp8","gpus":1,"source":"Gemma 4 31B: About 31 GB of FP8 weights (hiring-screen stack.json), or 18 GB at 4 bits (migration-check stack.json); the KV cache on top is an estimate."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"filing-preflight","name":"Filing pre-flight","page":"https://decosa.ai/legal/filing-preflight#self-host","tiers":[{"id":"lite","label":"Lite · CPU only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)","name":"decosa-api filing pre-flight (decosa_api/verticals/preflight)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the judge (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)","name":"decosa-api filing pre-flight (decosa_api/verticals/preflight)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per proposition and per record cite, plus one call for minors' names","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · a GLM-5.3-Flash judge on your own hardware","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)","name":"decosa-api filing pre-flight (decosa_api/verticals/preflight)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per proposition and per record cite, plus one call for minors' names","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-judge","role":"Stronger judge for propositions (wanted)","name":"GLM-5.3-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-one-card-judge","label":"A second judge on one card","gpu_gb":81.6,"basis":"estimate","unknown":[],"components":[{"id":"super-judge","role":"One-card second judge from another model family","name":"Nemotron-3-Super-120B-A12B (NVFP4)","where":"gpu","load":"resident","gb":81.6,"min_gb":81.6,"weights_gb":67.2,"basis":"estimate","precision":"fp4","gpus":1,"source":"Estimate: 120B parameters at about 4.5 bits per weight is about 67.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"promo-claims-check","name":"Promotional-claims pre-check","page":"https://decosa.ai/tools/life-sciences/promo-claims-check#self-host","tiers":[{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"precheck","role":"Pre-check: sentences, label sections, rules, findings, the packet (no model; CPU)","name":"decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Grounding judge, claim reviewer, head-to-head and fair-balance checks","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-vision","label":"Layout-aware check","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"vision","role":"Layout reader for visual prominence (alternate)","name":"Qwen3.8-27B vision input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"migration-check","name":"Open-model migration check","page":"https://decosa.ai/tools/developer/migration-check#self-host","tiers":[{"id":"lite","label":"Lite · structured prompts on a 24-32 GB card (self-host)","gpu_gb":38,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","name":"decosa-api migration module (decosa_api/verticals/migration)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"gemma","role":"Small candidate for structured tasks (self-host)","name":"Gemma-4-31B-it","where":"gpu","load":"resident","gb":38,"min_gb":34,"weights_gb":31,"basis":"estimate","precision":"fp8","gpus":1,"source":"Gemma 4 31B: About 31 GB of FP8 weights (hiring-screen stack.json), or 18 GB at 4 bits (migration-check stack.json); the KV cache on top is an estimate. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · Qwen3.8-27B as candidate and judge (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","name":"decosa-api migration module (decosa_api/verticals/migration)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Candidate and free-text judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash as the candidate (two 96 GB cards, self-host)","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","name":"decosa-api migration module (decosa_api/verticals/migration)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"deepseek","role":"Large candidate (self-host, two GPUs)","name":"DeepSeek-V4-Flash-0731","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate. (stack.json lists 176 GB for this component.)"}]},{"id":"wanted","label":"Wanted · the largest open candidates","gpu_gb":1377.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: rendering, scoring, statistics, cost, signed record (no model; CPU)","name":"decosa-api migration module (decosa_api/verticals/migration)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Candidate and free-text judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"dsv41-wanted","role":"Candidate: the largest open flash model","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."},{"id":"glm-wanted","role":"Candidate: the most-used open model","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"typed-judgment","name":"Typed-judgment API","page":"https://decosa.ai/tools/developer/typed-judgment#self-host","tiers":[{"id":"lite","label":"Lite · one call per question (stated confidence)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"engine","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","name":"decosa-api judgment module (decosa_api/verticals/judgment)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one temperature-0 call per question, plus seeded samples","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · hosted, answer plus 4 seeded samples","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"engine","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","name":"decosa-api judgment module (decosa_api/verticals/judgment)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one temperature-0 call per question, plus seeded samples","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"best","label":"Best · self-host with logprobs","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"engine","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","name":"decosa-api judgment module (decosa_api/verticals/judgment)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one temperature-0 call per question, plus seeded samples","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · two large judges that must agree","gpu_gb":1377.6,"basis":"estimate","unknown":[],"components":[{"id":"engine","role":"Engine: validation, prompts, parsing, calibration maps, eval mode, signed records (no model; CPU)","name":"decosa-api judgment module (decosa_api/verticals/judgment)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one temperature-0 call per question, plus seeded samples","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"dsv41-wanted","role":"Second judge for disagreement","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."},{"id":"glm-wanted","role":"Third judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-fast-judge","label":"Fast option-token judge","gpu_gb":57,"basis":"estimate","unknown":[],"components":[{"id":"gemma-judge","role":"Fast judge with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate. (stack.json lists 49 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.) The best tier fits too."},"rtx-5090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size). The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (32 of 96 GB). The best tier fits too."},"m5-max-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (32 of 64 GB). The best tier fits too."}}},{"id":"security-questionnaire","name":"Security questionnaire answerer","page":"https://decosa.ai/tools/finance/security-questionnaire#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card, no policy check","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"answerer","role":"Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)","name":"decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Selector (one call per question) and grounding judge (one call per changed or checked sentence)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one GPU, selection plus both checks (hosted demo)","gpu_gb":69.6,"basis":"stack","unknown":[],"components":[{"id":"answerer","role":"Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)","name":"decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"retrieval","role":"Candidate retrieval: the evidence retrieval block (dense + BM25, then a reranker) picks 4 approved answers and 2 policy passages per question","name":"Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","where":"gpu","load":"resident","gb":12,"min_gb":12,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 12 in stack.json."},{"id":"llm","role":"Selector (one call per question) and grounding judge (one call per changed or checked sentence)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · the largest open models, long context","gpu_gb":1377.6,"basis":"estimate","unknown":[],"components":[{"id":"answerer","role":"Answerer: parsing, candidate retrieval, copy checks, statuses, export and the signed record (no model; CPU)","name":"decosa-api questionnaire module (decosa_api/verticals/questionnaire), using the grounding module for checks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Selector (one call per question) and grounding judge (one call per changed or checked sentence)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"dsv41-wanted","role":"Selector and drafter with long context","name":"DeepSeek-V4.1-Flash","where":"gpu","load":"resident","gb":1128,"min_gb":800,"weights_gb":763,"basis":"stack","precision":"fp8","gpus":8,"source":"DeepSeek-V4.1-Flash: 763B parameters including Engram tables; the field stack puts it at about 8x H200 class (8 x 141 GB). The 800 GB minimum is an estimate."},{"id":"glm-wanted","role":"Checker, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-fast-judge","label":"Fast option-token selector","gpu_gb":57,"basis":"estimate","unknown":[],"components":[{"id":"gemma-selector","role":"Fast selector with option-token probabilities (alternate)","name":"Gemma-4-26B-A4B-it","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate. (stack.json lists 49 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 32 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (69.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (69.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"flight-recorder","name":"Agent flight recorder","page":"https://decosa.ai/tools/developer/flight-recorder#self-host","tiers":[{"id":"lite","label":"Lite · record only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"recorder","role":"Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)","name":"decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · receipted decisions, one 96 GB card (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"recorder","role":"Recorder: ingest API, hash chain, guards, sealing, thumbnails and verification (no model; runs on CPU)","name":"decosa-api flight recorder (decosa_api/verticals/flight) and the decosa_flight SDK","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"decider","role":"Decision model: picks the next action from a numbered element table (never coordinates, never free text)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"browser","role":"Headless browser for the hosted demo and the Playwright adapter","name":"Playwright 1.58 with Chromium headless shell","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-jev","label":"Typed decisions on Apple silicon (Jev-compatible)","gpu_gb":18.5,"basis":"stack","unknown":[],"components":[{"id":"jev","role":"Typed decisions on Apple silicon (Jev-compatible local decision server), for the macos-harness adapter","name":"DiffusionGemma 26B-A4B (MLX, OptiQ 4-bit)","where":"gpu","load":"resident","gb":18.5,"min_gb":18.5,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18.5 in stack.json."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"test-runs","name":"Verified end-to-end test runs","page":"https://decosa.ai/tools/developer/test-runs#self-host","tiers":[{"id":"lite","label":"Lite · navigation-only specs, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"runner","role":"Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)","name":"decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"browser","role":"Headless browser that carries out the steps and reads the assertions","name":"Playwright 1.58 with Chromium headless shell","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · agent steps, one 96 GB card (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Runner: spec parsing, the step loop, assertions in code, certificates, flake reports, domain proof (no model; runs on CPU)","name":"decosa-api test runs (decosa_api/verticals/testruns) on the flight recorder, and the decosa_testrun CI client","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"decider","role":"Action model: picks the next click, typing or selection from a numbered element table (never pass or fail)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"browser","role":"Headless browser that carries out the steps and reads the assertions","name":"Playwright 1.58 with Chromium headless shell","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-vision-actions","label":"Vision actions for canvas and iframe apps","gpu_gb":24.5,"basis":"estimate","unknown":[],"components":[{"id":"vision-actor","role":"Vision action model for canvas and iframe apps (controls not in the DOM)","name":"Holo-3.1-35B-A3B (NVFP4)","where":"gpu","load":"resident","gb":24.5,"min_gb":24.5,"weights_gb":19.6,"basis":"estimate","precision":"fp4","gpus":1,"source":"Estimate: 35B parameters at about 4.5 bits per weight is about 19.6 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"model-risk-pack","name":"Model-risk evidence pack","page":"https://decosa.ai/tools/finance/model-risk-pack#self-host","tiers":[{"id":"lite","label":"Lite · grader on one 32 GB card (self-host)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)","name":"decosa-api model-risk module (decosa_api/verticals/mrm)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Fixed grader (typed judgments) and the hosted system under test","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · hosted grader and demo systems (Qwen3.8-27B)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)","name":"decosa-api model-risk module (decosa_api/verticals/mrm)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Fixed grader (typed judgments) and the hosted system under test","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"clinical-ai-monitor","name":"Clinical AI assurance monitor","page":"https://decosa.ai/clinics/clinical-ai-monitor#self-host","tiers":[{"id":"lite","label":"Lite · text transcripts, one 32 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"monitor","role":"Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)","name":"decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"detail-checker","role":"Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5","name":"decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · judge plus diarizer (hosted demo)","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"monitor","role":"Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)","name":"decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"diarizer","role":"Audio in (optional): one-pass speaker-attributed transcript of the whole visit","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"detail-checker","role":"Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5","name":"decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"wanted","label":"Wanted · a second judge from another family","gpu_gb":253.6,"basis":"estimate","unknown":[],"components":[{"id":"monitor","role":"Monitor: transcript and note parsing, sentence and section offsets, finding types, the signed report and the summary with intervals and drift (no model; CPU)","name":"decosa-api clinical AI assurance monitor (decosa_api/verticals/monitor)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per note sentence, one checklist extraction, one coverage check; also the speaker role map for anonymous labels","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"judge-wanted","role":"Stronger judge (wanted)","name":"DeepSeek V4 Flash","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":166,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash (official FP8 + FP4 experts): About 83 GB of weights per GPU across two 96 GB cards at 32k context (clinical stack.json); little room left. The minimum is an estimate. (stack.json lists 166 GB for this component.)"},{"id":"diarizer","role":"Audio in (optional): one-pass speaker-attributed transcript of the whole visit","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"detail-checker","role":"Detail checker (M17): after the judge, each drug, dose, frequency, route, date, duration, side and number in a sentence is read against its transcript lines and labelled same / changed / absent; a judge 'detail not in the visit' flag becomes a changed-detail error when p(changed) >= 0.5","name":"decosa-note-detail-checker-modernbert-large (M17, our own model; Apache-2.0 on Hugging Face)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-one-card-judge","label":"A second judge on one card","gpu_gb":81.6,"basis":"estimate","unknown":[],"components":[{"id":"super-judge","role":"One-card second judge from another model family","name":"Nemotron-3-Super-120B-A12B (NVFP4)","where":"gpu","load":"resident","gb":81.6,"min_gb":81.6,"weights_gb":67.2,"basis":"estimate","precision":"fp4","gpus":1,"source":"Estimate: 120B parameters at about 4.5 bits per weight is about 67.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"privilege-log","name":"Privilege review and privilege log","page":"https://decosa.ai/legal/privilege-log#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card, self-hosted","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"reviewer","role":"Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)","name":"decosa-api privilege module (decosa_api/verticals/privilege)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"reviewer","role":"Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)","name":"decosa-api privilege module (decosa_api/verticals/privilege)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · a larger second judge on your own hardware","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"reviewer","role":"Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)","name":"decosa-api privilege module (decosa_api/verticals/privilege)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-wanted","role":"Second judge on the calls the 27B is least sure of","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-fast-judge","label":"Fast option-token judge for volume","gpu_gb":57,"basis":"estimate","unknown":[],"components":[{"id":"gemma-judge","role":"Faster model for the typed calls at volume (alternate)","name":"Gemma-4-26B-A4B-it","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate. (stack.json lists 49 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"editorial-ledger","name":"Editorial-control ledger","page":"https://decosa.ai/tools/media/editorial-ledger#self-host","tiers":[{"id":"lite","label":"Lite · ledger only, on CPU (self-host)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"ledger","role":"Ledger: edit metrics, sentence alignment, hash chain, sign-off rules, sealing and verification (no model; CPU)","name":"decosa-api editorial module (decosa_api/verticals/editorial)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Optional C2PA content credential on a DOCX copy of the published text","name":"c2pa-python 0.37 (native c2pa-rs)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · Qwen3.8-27B drafts and checks claims (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"ledger","role":"Ledger: edit metrics, sentence alignment, hash chain, sign-off rules, sealing and verification (no model; CPU)","name":"decosa-api editorial module (decosa_api/verticals/editorial)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Writes the AI first draft from the sources, and lists the claims an edit added or removed","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"c2pa","role":"Optional C2PA content credential on a DOCX copy of the published text","name":"c2pa-python 0.37 (native c2pa-rs)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"report-integrity","name":"Report integrity","page":"https://decosa.ai/legal/report-integrity#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":44,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same four steps on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Key events, first-draft writer, sentence judge (the grounding checker) and event coverage","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger judge for long, many-speaker incidents","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":388,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger judge for long, many-speaker incidents","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"legal-drafting-editor","name":"Privileged drafting editor","page":"https://decosa.ai/legal/legal-drafting-editor#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card, self-hosted","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"editor","role":"Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)","name":"decosa-api drafting module (decosa_api/verticals/drafting)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"reader","role":"Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clause","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"editor","role":"Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)","name":"decosa-api drafting module (decosa_api/verticals/drafting)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"reader","role":"Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clause","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · a larger second reader on your own hardware","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"editor","role":"Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)","name":"decosa-api drafting module (decosa_api/verticals/drafting)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"reader","role":"Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clause","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-wanted","role":"Second reader for clause detection","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-fast-judge","label":"Fast reader for long contracts","gpu_gb":57,"basis":"estimate","unknown":[],"components":[{"id":"gemma-reader","role":"Faster model for the clause-by-clause reading at volume (alternate)","name":"Gemma-4-26B-A4B-it","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate. (stack.json lists 49 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"foia-desk","name":"Public-records desk","page":"https://decosa.ai/legal/foia-desk#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card, self-hosted","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"desk","role":"Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU)","name":"decosa-api foia module (decosa_api/verticals/foia)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's calls","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"desk","role":"Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU)","name":"decosa-api foia module (decosa_api/verticals/foia)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's calls","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"oral-assessment","name":"Structured oral assessment","page":"https://decosa.ai/tools/operations/oral-assessment#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card, captions only","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers): the examiner prompt reads these","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm-lite","role":"Lite tier: the same scorer on the FP8 checkpoint; captions become the transcript (roles guessed)","name":"Qwen3.8-27B (official FP8)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."}]},{"id":"standard","label":"Standard · the hosted demo, two recognisers","gpu_gb":85.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers): the examiner prompt reads these","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After the session: examiner and candidate lines, each with a receipt over its audio","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm","role":"Examiner prompt, speaker roles, one score per criterion, grounding check of the evidence","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash scores and checks","gpu_gb":220,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions (streaming, no speakers): the examiner prompt reads these","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After the session: examiner and candidate lines, each with a receipt over its audio","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm-best","role":"Best tier: scorer and grounding judge","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]}],"mac":{"fit":"full","memory_gb":48,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 40 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 48 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (48 of 64 GB)."}}},{"id":"patent-claim-support","name":"Patent claim-support checker","page":"https://decosa.ai/legal/patent-claim-support#self-host","tiers":[{"id":"lite","label":"Lite · code checks only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)","name":"decosa-api patent module (decosa_api/verticals/patent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)","name":"decosa-api patent module (decosa_api/verticals/patent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: reads each claim element against the specification and names the supporting paragraphs","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · a GLM-5.3-Flash judge on your own hardware","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: claim parser, antecedent basis, claim references, claim terms, the claim chart, signed record and ledger (no model; CPU)","name":"decosa-api patent module (decosa_api/verticals/patent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: reads each claim element against the specification and names the supporting paragraphs","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-judge","role":"Stronger open model for the support map (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-one-card-judge","label":"A second judge on one card","gpu_gb":81.6,"basis":"estimate","unknown":[],"components":[{"id":"super-judge","role":"One-card second judge from another model family","name":"Nemotron-3-Super-120B-A12B (NVFP4)","where":"gpu","load":"resident","gb":81.6,"min_gb":81.6,"weights_gb":67.2,"basis":"estimate","precision":"fp4","gpus":1,"source":"Estimate: 120B parameters at about 4.5 bits per weight is about 67.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"music-gen-cleared","name":"Rights-cleared music generation","page":"https://decosa.ai/apps/music-gen-cleared#self-host","tiers":[{"id":"lite","label":"Lite · MIT engine, about 15 GB of GPU","gpu_gb":72.2,"basis":"stack","unknown":[],"components":[{"id":"guard","role":"Prompt guard, code layer: gazetteer of well-known names, look-alike folding, cover/cloning/type-beat patterns, a genre and era allow-list; the certificate and refusal log (CPU)","name":"decosa-api music module (decosa_api/verticals/music)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Prompt guard, typed judgment: names the category (artist, work, voice, label or none) with a calibrated probability; one call per brief","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"acestep","role":"Music generation: fast drafts and instrumentals (MIT engine)","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"similarity","role":"Similarity check (CPU): melody by chroma alignment across keys and tempos, sound by CLAP window embeddings, both normalised per reference","name":"LAION CLAP larger_clap_music + librosa chroma (services/music_embed)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, MiniMax-Music3 on a 96 GB card","gpu_gb":107.6,"basis":"stack","unknown":[],"components":[{"id":"guard","role":"Prompt guard, code layer: gazetteer of well-known names, look-alike folding, cover/cloning/type-beat patterns, a genre and era allow-list; the certificate and refusal log (CPU)","name":"decosa-api music module (decosa_api/verticals/music)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Prompt guard, typed judgment: names the category (artist, work, voice, label or none) with a calibrated probability; one call per brief","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"music3","role":"Music generation: songs with vocals and lyrics (default engine)","name":"MiniMax-Music3","where":"gpu","load":"job","gb":50,"min_gb":50,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-Music3: About 50 GB free on the card for Music3 (music-gen-cleared stack.json); the MLX port peaked at 50 GB on the Mac (mac.json)."},{"id":"acestep","role":"Music generation: fast drafts and instrumentals (MIT engine)","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"similarity","role":"Similarity check (CPU): melody by chroma alignment across keys and tempos, sound by CLAP window embeddings, both normalised per reference","name":"LAION CLAP larger_clap_music + librosa chroma (services/music_embed)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":{"fit":"partial","memory_gb":64,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 70 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 78 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 78 GB of GPU memory at the smallest settings; 64 GB available. The lite tier fits with changes."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 83.6 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 83.6 GB of GPU memory at the smallest settings; 80 GB available. The lite tier fits with changes."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (122.2 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: ACE-Step on a Mac is untested"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: ACE-Step on a Mac is untested"}}},{"id":"sample-clearance","name":"Sample and lyric clearance pre-check","page":"https://decosa.ai/tools/media/sample-clearance#self-host","tiers":[{"id":"lite","label":"Lite · code only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"matcher","role":"Audio matcher: log-frequency landmarks with pitch and tempo search, peak-by-peak verification, melody shingles for composition references; lyric spans; declared vs detected; the signed record (no model; CPU)","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · adds lyric labels by Qwen3.8-27B (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"matcher","role":"Audio matcher: log-frequency landmarks with pitch and tempo search, peak-by-peak verification, melody shingles for composition references; lyric spans; declared vs detected; the signed record (no model; CPU)","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"embed","role":"Lyric-line embeddings: re-worded lines that share few characters with the original","name":"all-MiniLM-L6-v2 (ONNX)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: labels each near-duplicate lyric line lift, variant, stock phrase or different","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"split-sheet-check","name":"Split-sheet and metadata checker","page":"https://decosa.ai/tools/media/split-sheet-check#self-host","tiers":[{"id":"lite","label":"Lite · code only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: DDEX and CSV parsing, OCR, identifier check digits, exact share arithmetic, grounding of every model-read value, reconciliation, the proposed sheet and signed records (no model; CPU)","name":"decosa-api splits module (decosa_api/verticals/splits) with Tesseract OCR","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: DDEX and CSV parsing, OCR, identifier check digits, exact share arithmetic, grounding of every model-read value, reconciliation, the proposed sheet and signed records (no model; CPU)","name":"decosa-api splits module (decosa_api/verticals/splits) with Tesseract OCR","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"reader","role":"Model: reads split sheets and contract excerpts (values copied as written, with line numbers) and answers typed same-writer and role questions","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-vl-reader","label":"A vision-language reader for scans","gpu_gb":20.9,"basis":"estimate","unknown":[],"components":[{"id":"vl-reader","role":"Vision-language model to read scanned or handwritten split sheets directly (alternate)","name":"Qwen2.5-VL-7B-Instruct","where":"gpu","load":"resident","gb":20.9,"min_gb":20.9,"weights_gb":16.6,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 8.3B parameters at 2 bytes (BF16 assumed; the quantisation is not stated) per weight is about 16.6 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"paper-claim-check","name":"Citation and claim checker for papers","page":"https://decosa.ai/tools/life-sciences/paper-claim-check#self-host","tiers":[{"id":"lite","label":"Lite · references only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)","name":"decosa-api papercheck module (decosa_api/verticals/papercheck)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"grobid","role":"Reference parser (optional): splits each reference into authors, title, year and DOI","name":"GROBID 0.8.2 (CRF models)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"checker","role":"Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)","name":"decosa-api papercheck module (decosa_api/verticals/papercheck)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"grobid","role":"Reference parser (optional): splits each reference into authors, title, year and DOI","name":"GROBID 0.8.2 (CRF models)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: reads each citing sentence against the cited paper's passages, then reviews its own flags","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"wanted","label":"Wanted · GLM-5.3-Flash for the claim check","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: manuscript and reference parsing, metadata lookups, retraction and citation-error checks, self-citation, signed report (no model; CPU)","name":"decosa-api papercheck module (decosa_api/verticals/papercheck)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"grobid","role":"Reference parser (optional): splits each reference into authors, title, year and DOI","name":"GROBID 0.8.2 (CRF models)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: reads each citing sentence against the cited paper's passages, then reviews its own flags","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-judge","role":"Stronger open model for the claim check (wanted)","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":{"fit":"partial","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested)"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: GROBID's image is amd64 only; on Apple Silicon it runs emulated (untested)"}}},{"id":"signed-lab-notebook","name":"Signed lab notebook","page":"https://decosa.ai/tools/life-sciences/signed-lab-notebook#self-host","tiers":[{"id":"lite","label":"Lite · notebook, signatures and timestamps on CPU (self-host)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"notebook","role":"Notebook: hash chain, members, amendments, e-signature rules, RFC 3161 client, export, audit CSV and verification (no model; CPU)","name":"decosa-api notebook module (decosa_api/verticals/notebook)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · Qwen3.8-27B writes receipted AI analyses (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"notebook","role":"Notebook: hash chain, members, amendments, e-signature rules, RFC 3161 client, export, audit CSV and verification (no model; CPU)","name":"decosa-api notebook module (decosa_api/verticals/notebook)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"qwen","role":"Writes the AI analysis note from the selected entries and attached data files","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"green-claims-check","name":"Green-claims substantiation check","page":"https://decosa.ai/tools/finance/green-claims-check#self-host","tiers":[{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"rulepack","role":"Rulepack, sentences, evidence spans, verdicts, claim table and signed report (no model; CPU)","name":"decosa-api green module (decosa_api/verticals/green) on the promo pre-check engine, with the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Claim typing, grounding judge, evidence reader and rewrite","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-vision","label":"Pack artwork and other languages","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"vision","role":"Pack artwork and label reader (alternate)","name":"Qwen3.8-27B vision input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"fi-disclosure-record","name":"Auto F&I disclosure record","page":"https://decosa.ai/tools/finance/fi-disclosure-record#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":44,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same extraction and checks on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Extraction (what was said about each add-on, prices, payment) and one typed yes/no/unclear check per question","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, many-speaker F&I sessions","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":388,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, many-speaker F&I sessions","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /fi/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"consent-ledger","name":"Consent ledger and replica gate","page":"https://decosa.ai/developers#block-consent-gate","tiers":[{"id":"lite","label":"Lite · the ledger and gate only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"ledger","role":"The ledger and gate: signed entries, one hash chain of events (enrol, revoke, suspend, strike), signed decisions, the C2PA link (no model; CPU)","name":"decosa-api consent module (decosa_api/verticals/consent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · adds speaker verification and scans on CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"ledger","role":"The ledger and gate: signed entries, one hash chain of events (enrol, revoke, suspend, strike), signed decisions, the C2PA link (no model; CPU)","name":"decosa-api consent module (decosa_api/verticals/consent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"speaker","role":"Speaker verification: a voice sample or a render against the enrolled consent clip, and the ingest scan of a finished cut","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"tts","role":"Demo renders only: the roster's stock voices speak a fixed line after the gate allows it","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · adds the use-description review by Qwen3.8-27B (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"ledger","role":"The ledger and gate: signed entries, one hash chain of events (enrol, revoke, suspend, strike), signed decisions, the C2PA link (no model; CPU)","name":"decosa-api consent module (decosa_api/verticals/consent)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"speaker","role":"Speaker verification: a voice sample or a render against the enrolled consent clip, and the ingest scan of a finished cut","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"tts","role":"Demo renders only: the roster's stock voices speak a fixed line after the gate allows it","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"review","role":"Advisory review: is each use description reasonably specific? One call per description","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (0 of 0 GB)."},"rtx-4090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 24 GB). The best tier fits too."},"rtx-5090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 32 GB). The best tier fits too."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 64 GB). The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 80 GB). The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 72 GB). The best tier fits too."},"m5-max-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."}}},{"id":"consented-dubbing","name":"Consented creator dubbing","page":"https://decosa.ai/apps/consented-dubbing#self-host","tiers":[{"id":"lite","label":"Lite · voice on CPU","gpu_gb":85.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Transcribes each speech span of the source, and listens back to every dubbed line","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewer","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"tts","role":"Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermark","name":"Chatterbox Multilingual (t3_23lang)","where":"gpu","load":"job","gb":4,"min_gb":4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4 in stack.json."},{"id":"verifier","role":"Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (use case 47)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, voice on a shared GPU","gpu_gb":85.6,"basis":"stack","unknown":[],"components":[{"id":"asr","role":"Transcribes each speech span of the source, and listens back to every dubbed line","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"llm","role":"Translation with the glossary and a length budget per line, repair of lines that break the glossary or budget, back-translation for the reviewer","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"tts","role":"Speaks each line in the consented voice (the enrolled consent clip is the reference), with the Perth watermark","name":"Chatterbox Multilingual (t3_23lang)","where":"gpu","load":"job","gb":4,"min_gb":4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4 in stack.json."},{"id":"verifier","role":"Consent ledger's speaker check: the reference clip must match the enrolled voiceprint (use case 47)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-voice-lipsync","label":"Better voice and lip-sync","gpu_gb":2.2,"basis":"estimate","unknown":["lipsync"],"components":[{"id":"tts-v3","role":"Alternate: the newer multilingual weights and the Latin American Spanish single-language pack from the same repository","name":"Chatterbox Multilingual V3 / Single Language Pack (LatAm Spanish)","where":"gpu","load":"job","gb":2.2,"min_gb":2.2,"weights_gb":1,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 0.5B parameters at 2 bytes (BF16 assumed; the quantisation is not stated) per weight is about 1 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"lipsync","role":"Alternate: lip-sync for talking-head video","name":"InfiniteTalk","where":"gpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Memory not stated in stack.json and not derivable (no parameter count)."}]}],"mac":null,"undetermined":[{"id":"lipsync","name":"InfiniteTalk","tiers":["alternate-voice-lipsync"]}],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 40 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 48 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Chatterbox Multilingual (t3_23lang) has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Chatterbox Multilingual (t3_23lang) has no mapped Apple Silicon build."}}},{"id":"disclosure-preflight","name":"Synthetic-performer disclosure and S&P pre-flight","page":"https://decosa.ai/tools/media/disclosure-preflight#self-host","tiers":[{"id":"lite","label":"Lite · code only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"preflight","role":"The pre-flight: provenance lookup, performer evidence, rules for NY and the EU, consent lookups by ledger id, the profanity and music-cue word lists, the report, sign-off and the signed attestation (no model; CPU)","name":"decosa-api disclosure module (decosa_api/verticals/disclosure)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ocr","role":"Reads six frames for an AI label (whole frame, then the top and bottom bands), before and after the label is added","name":"Tesseract OCR 5 (English)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Reads the file's C2PA credential and signs the marking: the source as a parentOf ingredient, c2pa.opened and c2pa.edited actions with the IPTC digital source type, and an ai.decosa.disclosure assertion","name":"c2pa-python 0.37 (c2pa-rs)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · adds script checks by Qwen3.8-27B (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"preflight","role":"The pre-flight: provenance lookup, performer evidence, rules for NY and the EU, consent lookups by ledger id, the profanity and music-cue word lists, the report, sign-off and the signed attestation (no model; CPU)","name":"decosa-api disclosure module (decosa_api/verticals/disclosure)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ocr","role":"Reads six frames for an AI label (whole frame, then the top and bottom bands), before and after the label is added","name":"Tesseract OCR 5 (English)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"Reads the file's C2PA credential and signs the marking: the source as a parentOf ingredient, c2pa.opened and c2pa.edited actions with the IPTC digital source type, and an ai.decosa.disclosure assertion","name":"c2pa-python 0.37 (c2pa-rs)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"watermark","role":"Decodes Decosa's invisible video watermark (TrustMark), so a Decosa render whose credential was stripped is still recognised","name":"TrustMark Q (decoder)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Model: lists candidate names, brands, profanity, rating triggers and music cues in the script, then answers a typed yes/no for each (typed-judgment, calibrated probability)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"animatic-studio","name":"Script to animatic","page":"https://decosa.ai/apps/animatic-studio#self-host","tiers":[{"id":"lite","label":"Lite · breakdown only, no GPU renderer","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Drafts the shot list (line ids, framing, action, duration) and a look per character; the same model is the grounding judge that checks each shot against its lines","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"standard","label":"Standard · the hosted demo","gpu_gb":91.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Drafts the shot list (line ids, framing, action, duration) and a look per character; the same model is the grounding judge that checks each shot against its lines","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"frames","role":"Character sheets, one still per shot (text only: every frame repeats each visible character's look), and optional image-to-video motion from a shot's still","name":"Wan2.2-VACE-Fun-A14B","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."},{"id":"lightning","role":"4-step distillation LoRAs for Wan2.2 (used at 6 steps, cfg 2 for stills)","name":"Wan2.2-Lightning T2V 4-step LoRAs","where":"gpu","load":"job","gb":0,"min_gb":0,"weights_gb":null,"basis":"estimate","precision":null,"gpus":1,"source":"A LoRA adapter loaded into the model it modifies; its memory is counted with that model (estimate: a few hundred MB)."},{"id":"tts","role":"Temp dialogue in stock voicepacks, only for speakers the consent ledger allows","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"C2PA content credential on the MP4, with the consent-ledger links of the voiced speakers","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · self-host with LTX-2.3 or MiniMax H3 (licence pending) motion","gpu_gb":153.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Drafts the shot list (line ids, framing, action, duration) and a look per character; the same model is the grounding judge that checks each shot against its lines","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"frames","role":"Character sheets, one still per shot (text only: every frame repeats each visible character's look), and optional image-to-video motion from a shot's still","name":"Wan2.2-VACE-Fun-A14B","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."},{"id":"lightning","role":"4-step distillation LoRAs for Wan2.2 (used at 6 steps, cfg 2 for stills)","name":"Wan2.2-Lightning T2V 4-step LoRAs","where":"gpu","load":"job","gb":0,"min_gb":0,"weights_gb":null,"basis":"estimate","precision":null,"gpus":1,"source":"A LoRA adapter loaded into the model it modifies; its memory is counted with that model (estimate: a few hundred MB)."},{"id":"tts","role":"Temp dialogue in stock voicepacks, only for speakers the consent ledger allows","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"c2pa","role":"C2PA content credential on the MP4, with the consent-ledger links of the voiced speakers","name":"c2pa-rs via c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ltx","role":"Longer, higher-resolution motion shots on your own GPU","name":"LTX-2.3 22B","where":"gpu","load":"job","gb":50,"min_gb":44,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"LTX-2 22B video: LTX-2.5 distilled renders within about 50 GB of free VRAM (measured 2026-09-23, studio stack.json); LTX-2.3 with a character LoRA about 44 GB (owner's pipeline notes, characters stack.json)."},{"id":"h3","role":"Motion shots with native audio from a shot's still (image-to-video)","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."}]},{"id":"alternate-ref-edit","label":"Reference-based consistency","gpu_gb":49,"basis":"estimate","unknown":[],"components":[{"id":"ref-edit","role":"Reference-image editing to keep a character's face and costume from the sheet","name":"Qwen-Image-Edit","where":"gpu","load":"resident","gb":49,"min_gb":49,"weights_gb":40,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 20B parameters at 2 bytes (BF16 assumed; the quantisation is not stated) per weight is about 40 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 54 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 62 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Wan2.2-VACE-Fun-A14B needs about 34 GB on one GPU; each GPU here has 32 GB. The lite tier fits with changes."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 67.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (91.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (91.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"audio-drama-studio","name":"Audio drama and narrated story studio","page":"https://decosa.ai/apps/audio-drama-studio#self-host","tiers":[{"id":"lite","label":"Lite · CPU only, radio scripts","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"tts","role":"Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPU","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"mixer","role":"Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)","name":"decosa-api drama module (decosa_api/verticals/drama) + FFmpeg","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo","gpu_gb":72.2,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"tts","role":"Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPU","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"music","role":"Score: theme and sting cues rendered through the music-gen-cleared path (use case 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPU","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"mixer","role":"Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)","name":"decosa-api drama module (decosa_api/verticals/drama) + FFmpeg","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"verifier","role":"Consent ledger's speaker check (use case 47): does each role's rendered voice match the voice enrolled in its entry?","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · compose new music per episode","gpu_gb":72.2,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Parse: voice hints, aliases and the sound and music cue mapping for radio scripts (one call); speaker attribution and sound suggestions for prose (one call per 36 quotations)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"tts","role":"Speaks each line with a stock voicepack after the consent ledger allows it; one process per episode on CPU","name":"Kokoro-82M","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"music","role":"Score: theme and sting cues rendered through the music-gen-cleared path (use case 37) with its prompt guard, similarity check and signed licence certificate; the hosted demo uses the library, composing new cues needs the studio GPU","name":"ACE-Step 1.5 turbo + 5Hz LM 1.7B","where":"gpu","load":"job","gb":14.6,"min_gb":14.6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 14.6 in stack.json."},{"id":"mixer","role":"Timeline, sound library, ducking, compression, limiter, loudness to spec, chapters, captions, sides, C2PA and the signed record (CPU)","name":"decosa-api drama module (decosa_api/verticals/drama) + FFmpeg","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"verifier","role":"Consent ledger's speaker check (use case 47): does each role's rendered voice match the voice enrolled in its entry?","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-expressive-voices","label":"Voices that can act","gpu_gb":4,"basis":"stack","unknown":[],"components":[{"id":"expressive-tts","role":"Alternate: a permissively licensed voice model with emotion control, driven by stock (not cloned) reference voices","name":"Chatterbox (exaggeration control)","where":"gpu","load":"job","gb":4,"min_gb":4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 34.6 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 42.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 48.2 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (72.2 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (72.2 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: ACE-Step 1.5 turbo + 5Hz LM 1.7B has no mapped Apple Silicon build The lite tier fits."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: ACE-Step 1.5 turbo + 5Hz LM 1.7B has no mapped Apple Silicon build The lite tier fits."}}},{"id":"music-video-studio","name":"Music video from your track","page":"https://decosa.ai/apps/music-video-studio#self-host","tiers":[{"id":"lite","label":"Lite · timed captions and a cited treatment, no GPU for video","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"analyzer","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"aligner","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","name":"wav2vec2-large-960h-lv60-self","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"clearance","role":"Sample and lyric clearance pre-check on the upload (use case 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, clips on one 96 GB card","gpu_gb":91.6,"basis":"stack","unknown":[],"components":[{"id":"analyzer","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"aligner","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","name":"wav2vec2-large-960h-lv60-self","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"clearance","role":"Sample and lyric clearance pre-check on the upload (use case 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"video","role":"Clips: text-to-video, one 5 s clip per scene per format","name":"Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."},{"id":"edit","role":"The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (use case 49) on the file; C2PA credential per file; the signed record","name":"decosa-api mvideo module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · MiniMax H3 clips (licence pending), self-host","gpu_gb":153.6,"basis":"stack","unknown":[],"components":[{"id":"analyzer","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"aligner","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","name":"wav2vec2-large-960h-lv60-self","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"clearance","role":"Sample and lyric clearance pre-check on the upload (use case 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"h3","role":"A clip per scene with native audio, in place of Wan2.2","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."},{"id":"edit","role":"The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (use case 49) on the file; C2PA credential per file; the signed record","name":"decosa-api mvideo module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"wanted","label":"Wanted · MiniMax H3 on two cards, no offload","gpu_gb":153.6,"basis":"stack","unknown":[],"components":[{"id":"analyzer","role":"Beat, bar and section detection (CPU): librosa beat tracker on a full-band plus low-band onset envelope, bar phase by a kick-and-snare heuristic (4/4), sections by checkerboard novelty on chroma and MFCC self-similarity","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"aligner","role":"Lyric timing (CPU): CTC forced alignment of the artist's own lyrics on the mix, with a repair pass for lines squeezed into too little time; English letters","name":"wav2vec2-large-960h-lv60-self","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Treatment writer (scenes that cite the lyric lines they show) and the typed yes/no grounding check per scene; also labels near-duplicate lyric lines in the clearance pre-check","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"clearance","role":"Sample and lyric clearance pre-check on the upload (use case 38, run in-process): audio landmarks and melody against a small open catalogue, lyric lines against a lyric set","name":"decosa-api clearance module (decosa_api/verticals/clearance)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"h3","role":"A clip per scene with native audio, in place of Wan2.2","name":"MiniMax-H3 (licence pending)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json)."},{"id":"edit","role":"The edit and the marks (CPU): cuts on bar lines on a 30 fps grid, karaoke captions (ASS, libass), the AI label on a top bar, the credit, the artist's audio; cut timing measured back from the pixels; the disclosure pre-flight's rules (use case 49) on the file; C2PA credential per file; the signed record","name":"decosa-api mvideo module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-ltx","label":"Sharper clips with LTX-2.3","gpu_gb":50,"basis":"stack","unknown":[],"components":[{"id":"ltx","role":"Alternate, self-host only: sharper clips","name":"LTX-2.3 22B (distilled)","where":"gpu","load":"job","gb":50,"min_gb":44,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"LTX-2 22B video: LTX-2.5 distilled renders within about 50 GB of free VRAM (measured 2026-09-23, studio stack.json); LTX-2.3 with a character LoRA about 44 GB (owner's pipeline notes, characters stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 54 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 62 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs needs about 34 GB on one GPU; each GPU here has 32 GB. The lite tier fits with changes."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 67.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (91.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (91.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B + Wan2.2-Lightning 4-step LoRAs has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"sar-narrative-desk","name":"SAR narrative desk","page":"https://decosa.ai/tools/finance/sar-narrative-desk#self-host","tiers":[{"id":"lite","label":"Lite · numbers and citations only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"desk","role":"Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)","name":"decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"desk","role":"Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)","name":"decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Section drafting and the grounding judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-ocr","label":"Statements as PDFs and scans","gpu_gb":0,"basis":null,"unknown":["ocr"],"components":[{"id":"ocr","role":"Bank-statement and wire-confirmation reader (alternate)","name":"A licence-clean vision/OCR model (not chosen)","where":"none","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Not built or not chosen yet."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[{"id":"ocr","name":"A licence-clean vision/OCR model (not chosen)","tiers":["alternate-ocr"]}],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"claims-conduct-pack","name":"Insurance claims-file conduct pack","page":"https://decosa.ai/tools/finance/claims-conduct-pack#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same extraction, grounding and judgments on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the file (dated events, denial reasons, amounts, AI mentions, each with a quote), grounds each denial reason in the policy, and answers the typed conduct questions","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, multi-claimant files","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, multi-claimant files","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"collections-call-qa","name":"Collections and servicing call QA","page":"https://decosa.ai/tools/finance/collections-call-qa#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":44,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same extraction and checks on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Extraction (what the called person asked for or said, with the line) and one typed yes/no/unclear check per QA question","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long calls and hard cases","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":388,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long calls and hard cases","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /collections/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"incident-notification-pack","name":"Incident notification pack","page":"https://decosa.ai/tools/finance/incident-notification-pack#self-host","tiers":[{"id":"lite","label":"Lite · clocks and timeline, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"pack","role":"Timeline, hash chain, deadline clocks, citation, time and number checks, element coverage, cross-notice consistency, signed pack and sign-off (no model; CPU)","name":"decosa-api incident pack (decosa_api/verticals/incident) with the numeric block (decosa_api/verticals/numeric), the dates module (decosa_api/verticals/claims/dates.py) and the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"pack","role":"Timeline, hash chain, deadline clocks, citation, time and number checks, element coverage, cross-notice consistency, signed pack and sign-off (no model; CPU)","name":"decosa-api incident pack (decosa_api/verticals/incident) with the numeric block (decosa_api/verticals/numeric), the dates module (decosa_api/verticals/claims/dates.py) and the grounding module (decosa_api/verticals/grounding)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Notice drafting, the grounding judge, and the review call (element coverage and quoted facts)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"filing-tieout","name":"Filing tie-out and MD&A grounding","page":"https://decosa.ai/tools/finance/filing-tieout#self-host","tiers":[{"id":"lite","label":"Lite · figures only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"tieout","role":"iXBRL parser, figure tie-out, period and scale rules, direction and cross-reference checks, workpaper and signed record (no model; CPU)","name":"decosa-api tie-out (decosa_api/verticals/tieout) with the numeric grounding block's table-cell matcher (decosa_api/verticals/numeric/cells.py)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for claims in words (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"tieout","role":"iXBRL parser, figure tie-out, period and scale rules, direction and cross-reference checks, workpaper and signed record (no model; CPU)","name":"decosa-api tie-out (decosa_api/verticals/tieout) with the numeric grounding block's table-cell matcher (decosa_api/verticals/numeric/cells.py)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Claims in words (optional): the grounding judge reads a sentence against the named lines' figures","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-docreader","label":"MD&A tables and Word/PDF drafts","gpu_gb":0,"basis":null,"unknown":["docreader"],"components":[{"id":"docreader","role":"Tables inside the MD&A, and Word or PDF drafts (alternate)","name":"A licence-clean table and document reader (not chosen)","where":"none","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Not built or not chosen yet."}]}],"mac":null,"undetermined":[{"id":"docreader","name":"A licence-clean table and document reader (not chosen)","tiers":["alternate-docreader"]}],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"sanctions-disposition-record","name":"Sanctions alert disposition record","page":"https://decosa.ai/tools/finance/sanctions-disposition-record#self-host","tiers":[{"id":"lite","label":"Lite · the comparison and the proposal, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"desk","role":"List snapshot, field comparison, proposal rules, rationale check, signed record, audit sample (no model; CPU)","name":"decosa-api sanctions desk (decosa_api/verticals/sanctions), with the typed-judgment core (vertical 24) and the session hash chain (record, vertical 07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"desk","role":"List snapshot, field comparison, proposal rules, rationale check, signed record, audit sample (no model; CPU)","name":"decosa-api sanctions desk (decosa_api/verticals/sanctions), with the typed-judgment core (vertical 24) and the session hash chain (record, vertical 07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Independent typed reading and the rationale","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-translit","label":"Names in every script","gpu_gb":0,"basis":null,"unknown":["translit"],"components":[{"id":"translit","role":"Transliteration for Arabic, Persian, Chinese and Korean names (alternate)","name":"A licence-clean multilingual transliteration or name-matching model (not chosen)","where":"none","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Not built or not chosen yet."}]}],"mac":null,"undetermined":[{"id":"translit","name":"A licence-clean multilingual transliteration or name-matching model (not chosen)","tiers":["alternate-translit"]}],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"medicare-call-record","name":"Medicare sales-call record","page":"https://decosa.ai/clinics/medicare-call-record#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":44,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same extractions, checks and claim judge on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Two extractions (products, first benefit discussion, enrollment steps; the agent's benefit claims), one typed yes/no/unclear check per question, and one grounding judgment per benefit claim against the plan facts","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"best","label":"Best · two 96 GB cards","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long calls and hard cases","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":388,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long calls and hard cases","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /medicare/transcribe, self-host; the demo's audio samples were transcribed with it)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"denial-appeal-packet","name":"Claim denial appeal packet","page":"https://decosa.ai/clinics/denial-appeal-packet#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the denial (notice date, reasons, service), splits the payer policy into criteria, checks each criterion against the chart, drafts the letter when the chart supports it, and judges every letter sentence (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long charts and dense payer policies","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long charts and dense payer policies","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"kids-content-preflight","name":"Made-for-kids content pre-flight","page":"https://decosa.ai/tools/media/kids-content-preflight#self-host","tiers":[{"id":"lite","label":"Lite · audience only (the catalogue audit)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Evidence lines per FTC factor (quotes), a typed audience choice with probabilities, a typed yes/no per factor (typed-judgment, vertical 24); with vision on, a description of four frames","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"ocr","role":"Reads on-screen text from six sampled frames of a video","name":"Tesseract OCR 5 (English)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"preflight","role":"The pre-flight: intake, quote checks, word-list cues, the setting check, the kids-directed flags, the batch audit, sign-off and the signed review record (no model; CPU)","name":"decosa-api kids module (decosa_api/verticals/kids)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · full check on text (hosted demo)","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Evidence lines per FTC factor (quotes), a typed audience choice with probabilities, a typed yes/no per factor (typed-judgment, vertical 24); with vision on, a description of four frames","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Speech to timed lines for a video sent without a transcript, with a speech receipt signed by the instance","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"ocr","role":"Reads on-screen text from six sampled frames of a video","name":"Tesseract OCR 5 (English)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"preflight","role":"The pre-flight: intake, quote checks, word-list cues, the setting check, the kids-directed flags, the batch audit, sign-off and the signed review record (no model; CPU)","name":"decosa-api kids module (decosa_api/verticals/kids)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · adds the frame check (vision)","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Evidence lines per FTC factor (quotes), a typed audience choice with probabilities, a typed yes/no per factor (typed-judgment, vertical 24); with vision on, a description of four frames","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Speech to timed lines for a video sent without a transcript, with a speech receipt signed by the instance","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"ocr","role":"Reads on-screen text from six sampled frames of a video","name":"Tesseract OCR 5 (English)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"preflight","role":"The pre-flight: intake, quote checks, word-list cues, the setting check, the kids-directed flags, the batch audit, sign-off and the signed review record (no model; CPU)","name":"decosa-api kids module (decosa_api/verticals/kids)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.) The best tier fits too."},"rtx-5090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (61.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon. The best tier fits too."},"m5-max-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon. The best tier fits too."}}},{"id":"hcc-evidence-file","name":"HCC evidence file and RADV defence","page":"https://decosa.ai/clinics/hcc-evidence-file#self-host","tiers":[{"id":"lite","label":"Lite · map and record checks only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"review","role":"V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)","name":"decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"review","role":"V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)","name":"decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"best","label":"Best · scanned charts too (adds the document reader)","gpu_gb":63,"basis":"stack","unknown":[],"components":[{"id":"review","role":"V28 map and hierarchies, RADV record checks, verdict rules, the no-add guard, net effect, evidence file and signed record (no model; CPU)","name":"decosa-api HCC review (decosa_api/verticals/hcc), importing the grounding judge (vertical 17) and the typed-judgment engine (vertical 24)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Evidence sentences with MEAT tags, the typed judgment per record, and the grounding judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"layout","role":"Scanned charts: finds and orders the regions of each page (layout only)","name":"Docling 2.130 with the Heron layout model","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Scanned charts: reads each region (text, tables as cells); Qwen3.8 re-reads what it is unsure of and reads the header and signature fields","name":"PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"vex-triage","name":"Signed VEX triage","page":"https://decosa.ai/tools/developer/vex-triage#self-host","tiers":[{"id":"lite","label":"Lite · rules only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checks","role":"The collector (on your machine) and the checks in decosa-api: package in the SBOM, version against the fix and the OSV/NVD ranges, the Debian changelog, ELF loader chains, the advisory's programs and settings, the platform; the evidence gates; OpenVEX, CycloneDX VEX and DSSE signing.","name":"decosa VEX checks and collector (code, no model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"One typed judgment per finding the rules do not settle: a VEX status with quoted evidence from the advisory lines and the checks, and the impact statement. Code gates every answer.","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"checks","role":"The collector (on your machine) and the checks in decosa-api: package in the SBOM, version against the fix and the OSV/NVD ranges, the Debian changelog, ELF loader chains, the advisory's programs and settings, the platform; the evidence gates; OpenVEX, CycloneDX VEX and DSSE signing.","name":"decosa VEX checks and collector (code, no model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"device-mdr-triage","name":"Device complaint MDR triage","page":"https://decosa.ai/tools/life-sciences/device-mdr-triage#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the complaint (dates an employee was told, the event date, the device problem), answers the 803.50(a) questions one at a time (outcome, caused or contributed, malfunction, likely if it recurred), drafts the 3500A event description, judges every narrative sentence (the grounding judge), and groups complaints by failure mode for the trend view","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, messy complaint files","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, messy complaint files","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"cmmc-evidence-map","name":"CMMC / NIST 800-171 evidence map","page":"https://decosa.ai/tools/finance/cmmc-evidence-map#self-host","tiers":[{"id":"lite","label":"Lite · catalog and artefact checks only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"map","role":"Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU)","name":"decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"map","role":"Catalog, point values and POA&M rules, artefact checks, CUI guard, status rules, gap map, POA&M draft, evidence index and signed record (no model; CPU)","name":"decosa-api CMMC evidence map (decosa_api/verticals/cmmc), importing the grounding judge (17), the typed-judgment engine (24) and the test-run certificate verifier (27)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Screenshot transcription, lines per objective, the typed judgment per objective, and the grounding judge","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"alternate-ocr-lane","label":"Evidence binders as PDFs (document reader)","gpu_gb":0,"basis":null,"unknown":["ocr"],"components":[{"id":"ocr","role":"Scanned PDF and evidence-binder reader (wanted)","name":"A licence-clean document OCR and layout model (not chosen)","where":"none","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Not built or not chosen yet."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[{"id":"ocr","name":"A licence-clean document OCR and layout model (not chosen)","tiers":["alternate-ocr-lane"]}],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"eu-trial-lay-summary","name":"EU trial lay summary with number grounding","page":"https://decosa.ai/tools/life-sciences/eu-trial-lay-summary#self-host","tiers":[{"id":"lite","label":"Lite · check a draft in code, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"laysummary","role":"Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)","name":"decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":75.6,"basis":"stack","unknown":[],"components":[{"id":"laysummary","role":"Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)","name":"decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sources","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"translate","role":"Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech","name":"Hy-MT2-7B (the language-pack block)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."}]}],"mac":{"fit":"partial","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 38 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 46 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 51.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (75.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (75.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before."}}},{"id":"reg-e-dispute-file","name":"Reg E dispute investigation file","page":"https://decosa.ai/tools/finance/reg-e-dispute-file#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the consumer's notice (the required elements, the error asserted, who made the transfer, the transfers named), the investigator's notes (evidence reviewed, waiting for paperwork, carelessness cited) and the results letter (explanation specific or generic, right to documents, debit notice), and judges each letter sentence against the file (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long call transcripts and dense case notes","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long call transcripts and dense case notes","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"virtual-staging","name":"Disclosed virtual staging","page":"https://decosa.ai/tools/operations/virtual-staging#self-host","tiers":[{"id":"lite","label":"Lite · check only (any staging tool)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Maps the original's windows, doors, fixtures, damage and open floor; boxes each piece of furniture on the render; compares original and staged side by side; the typed judgment on instructions (typed-judgment, vertical 24)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"check","role":"The zone, the furniture-only composite, the pixel check (added, removed, recoloured regions; floor pattern; shadows; what each touches), the verdict, the burned-in label and its OCR read-back, the C2PA credential, the public original page and the signed pair record (no model; CPU)","name":"decosa-api staging module (decosa_api/verticals/staging)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · stage and check (hosted demo)","gpu_gb":91.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Maps the original's windows, doors, fixtures, damage and open floor; boxes each piece of furniture on the render; compares original and staged side by side; the typed judgment on instructions (typed-judgment, vertical 24)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"render","role":"Draws furniture inside the zone (the open floor, with windows, doors, radiators, built-ins and damage cut out): one inpainted still per attempt, VACE control image plus mask","name":"Wan2.2-VACE-Fun-A14B","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."},{"id":"lightning","role":"4-step distillation LoRAs for the render","name":"Wan2.2-Lightning T2V 4-step LoRAs","where":"gpu","load":"job","gb":0,"min_gb":0,"weights_gb":null,"basis":"estimate","precision":null,"gpus":1,"source":"A LoRA adapter loaded into the model it modifies; its memory is counted with that model (estimate: a few hundred MB)."},{"id":"check","role":"The zone, the furniture-only composite, the pixel check (added, removed, recoloured regions; floor pattern; shadows; what each touches), the verdict, the burned-in label and its OCR read-back, the C2PA credential, the public original page and the signed pair record (no model; CPU)","name":"decosa-api staging module (decosa_api/verticals/staging)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-upgrade","label":"instruction image editing","gpu_gb":49,"basis":"estimate","unknown":[],"components":[{"id":"edit","role":"Instruction-following image editing ('add a grey sofa by the left wall') that keeps the rest of the photo by design","name":"Qwen-Image-Edit-2511","where":"gpu","load":"job","gb":49,"min_gb":49,"weights_gb":40,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 20B parameters at 2 bytes (BF16 assumed; the quantisation is not stated) per weight is about 40 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 54 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 62 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Wan2.2-VACE-Fun-A14B needs about 34 GB on one GPU; each GPU here has 32 GB. The lite tier fits with changes."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 67.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (91.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (91.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Wan2.2-VACE-Fun-A14B has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"ma-dd-redflags","name":"M&A due-diligence red flags","page":"https://decosa.ai/legal/ma-dd-redflags#self-host","tiers":[{"id":"lite","label":"Lite · BM25 retrieval, no retrieval service","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"review","role":"Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)","name":"decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each finding","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · retrieval service + the model (hosted demo)","gpu_gb":66.6,"basis":"stack","unknown":[],"components":[{"id":"review","role":"Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)","name":"decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each finding","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"retrieval","role":"Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per search","name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","where":"gpu","load":"resident","gb":9,"min_gb":9,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 9 in stack.json."}]},{"id":"alternate-scans","label":"Scanned data rooms (document reader)","gpu_gb":66.6,"basis":"stack","unknown":[],"components":[{"id":"review","role":"Categories and queries, the quote gate, merging, the memo and the signed record (no model; CPU)","name":"decosa-api DD red flags (decosa_api/verticals/madd), importing the evidence retrieval block and the grounding judge (17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per category (which passages show a flag, in their exact words, why it matters, what to ask) and the grounding judge on each finding","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"retrieval","role":"Chunking with byte offsets, the index hash, hybrid search (dense + BM25) and reranking, with a signed receipt per search","name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","where":"gpu","load":"resident","gb":9,"min_gb":9,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 9 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 29 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 37 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (66.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (66.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"tariff-classification","name":"Tariff classification memo","page":"https://decosa.ai/tools/finance/tariff-classification#self-host","tiers":[{"id":"lite","label":"Lite · smaller reranker","gpu_gb":62.4,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"One call per memo: reads the rulings found, the HTS text of the candidate headings and the GRI, and proposes a heading, subheading and statistical number with quotes, or declines.","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"embedder","role":"Embeds the ruling set once (cached) and each product description, for the dense half of the hybrid search (BM25 is the other half).","name":"Qwen3-Embedding-0.6B","where":"gpu","load":"resident","gb":2.4,"min_gb":2.4,"weights_gb":1.2,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 0.6B parameters at 2 bytes (BF16) per weight is about 1.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"reranker-small","role":"The smaller reranker for a card with less memory.","name":"Qwen3-Reranker-0.6B","where":"gpu","load":"resident","gb":2.4,"min_gb":2.4,"weights_gb":1.2,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 0.6B parameters at 2 bytes (BF16) per weight is about 1.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"checks","role":"Codes against the HTS release, codes nest, quotes word for word with byte offsets, a cited ruling at the proposed heading, rulings agree, evidence not thin; the signed record.","name":"Checks and record (decosa-api, Python)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo","gpu_gb":70.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"One call per memo: reads the rulings found, the HTS text of the candidate headings and the GRI, and proposes a heading, subheading and statistical number with quotes, or declines.","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"embedder","role":"Embeds the ruling set once (cached) and each product description, for the dense half of the hybrid search (BM25 is the other half).","name":"Qwen3-Embedding-0.6B","where":"gpu","load":"resident","gb":2.4,"min_gb":2.4,"weights_gb":1.2,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 0.6B parameters at 2 bytes (BF16) per weight is about 1.2 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"reranker","role":"Scores the 40 best chunks for each description, so the rulings the model reads are the closest products, not the closest words.","name":"Qwen3-Reranker-4B","where":"gpu","load":"resident","gb":10.6,"min_gb":10.6,"weights_gb":8,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 4B parameters at 2 bytes (BF16) per weight is about 8 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"checks","role":"Codes against the HTS release, codes nest, quotes word for word with byte offsets, a cited ruling at the proposed heading, rulings agree, evidence not thin; the signed record.","name":"Checks and record (decosa-api, Python)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 33 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 41 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (70.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (70.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-Embedding-0.6B has no mapped Apple Silicon build; Qwen3-Reranker-4B has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-Embedding-0.6B has no mapped Apple Silicon build; Qwen3-Reranker-4B has no mapped Apple Silicon build."}}},{"id":"csr-number-verifier","name":"CSR number-to-table verifier","page":"https://decosa.ai/tools/life-sciences/csr-number-verifier#self-host","tiers":[{"id":"lite","label":"Lite · text and rows, keyword search","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"verifier","role":"Finds the numbers, keeps the tables as typed cells, compares and recomputes in code, names the likely slip, checks in-text tables against their TLF, writes the QC report and the signed record (no model; CPU)","name":"decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per paragraph: which cell each number claims to report (by arm, row and timepoint), or which cells a derived number is computed from; a second look for numbers left unplaced; a second reading (values hidden) between two neighbouring cells","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · retrieval + document reader + the model (hosted demo)","gpu_gb":72.6,"basis":"stack","unknown":[],"components":[{"id":"verifier","role":"Finds the numbers, keeps the tables as typed cells, compares and recomputes in code, names the likely slip, checks in-text tables against their TLF, writes the QC report and the signed record (no model; CPU)","name":"decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per paragraph: which cell each number claims to report (by arm, row and timepoint), or which cells a derived number is computed from; a second look for numbers left unplaced; a second reading (values hidden) between two neighbouring cells","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"retrieval","role":"Finds the tables a paragraph most likely reports in a long report (after the tables it names), with a signed receipt per search","name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","where":"gpu","load":"resident","gb":9,"min_gb":9,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 9 in stack.json."},{"id":"reader","role":"PDFs and scans: finds and orders the regions of each page (Docling Heron layout) and reads tables as cells with spans (PaddleOCR-VL-1.6); born-digital text comes from the PDF's text layer","name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."}]},{"id":"alternate-number-checker","label":"Number-consistency checker tier (CPU)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"number-checker","role":"Optional tier for numbers the pointer leaves untraced: a small encoder trained on planted number errors reads the sentence against the table","name":"Number-consistency checker (decosa_api.checkers.numbers, on a pre-release build)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 35 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 43 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 48.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (72.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (72.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"gpsr-listing-pack","name":"GPSR listing pack","page":"https://decosa.ai/tools/finance/gpsr-listing-pack#self-host","tiers":[{"id":"standard","label":"Standard · Qwen3.8-27B plus Hy-MT2-7B (hosted demo)","gpu_gb":75.6,"basis":"stack","unknown":[],"components":[{"id":"gpsr","role":"Article 19 checklist, value lookup against the documents, label-vs-sheet comparison, listing export and signed record (no model; CPU)","name":"decosa-api GPSR pack (decosa_api/verticals/gpsr), on the language-pack block (decosa_api/lang) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Label reading (image input) and field extraction as JSON","name":"Qwen3.8-27B (NVFP4, vision tower on)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"mt","role":"Translation of warnings and safety information (the language-pack block): German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech","name":"Hy-MT2-7B","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."}]},{"id":"best","label":"Best · the 30B translation model (not measured)","gpu_gb":94.6,"basis":"estimate","unknown":[],"components":[{"id":"gpsr","role":"Article 19 checklist, value lookup against the documents, label-vs-sheet comparison, listing export and signed record (no model; CPU)","name":"decosa-api GPSR pack (decosa_api/verticals/gpsr), on the language-pack block (decosa_api/lang) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Label reading (image input) and field extraction as JSON","name":"Qwen3.8-27B (NVFP4, vision tower on)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"mt-30b","role":"Translation, larger model","name":"Hy-MT2-30B-A3B-FP8","where":"gpu","load":"resident","gb":37,"min_gb":37,"weights_gb":30,"basis":"estimate","precision":"fp8","gpus":1,"source":"Estimate: 30B parameters at 1 byte per weight is about 30 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4, vision tower on) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 38 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 46 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4, vision tower on): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 51.6 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4, vision tower on) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (75.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (75.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Hy-MT2-7B has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Hy-MT2-7B has no mapped Apple Silicon build."}}},{"id":"medical-chronology","name":"Medical chronology with page cites","page":"https://decosa.ai/legal/medical-chronology#self-host","tiers":[{"id":"lite","label":"Lite · document reader and the model, no retrieval service","gpu_gb":63,"basis":"stack","unknown":[],"components":[{"id":"chronology","role":"The event schema and page checks, copies, merging, conflicts, gaps, pre-existing, cites to chunks and bytes, the Markdown chronology and the signed record (no model; CPU)","name":"decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per page (the page's elements, plus the page image for a scan) listing typed events with quote and date; typed judgments for duplicate and conflict pairs and for pre-existing conditions; re-reads of regions the parser was unsure of","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"layout","role":"Finds and orders the regions of every page (text, headings, tables, checkboxes) with their boxes","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned page (text, handwriting, tables as cells)","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]},{"id":"standard","label":"Standard · reader, retrieval service and the model (hosted demo)","gpu_gb":72,"basis":"stack","unknown":[],"components":[{"id":"chronology","role":"The event schema and page checks, copies, merging, conflicts, gaps, pre-existing, cites to chunks and bytes, the Markdown chronology and the signed record (no model; CPU)","name":"decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per page (the page's elements, plus the page image for a scan) listing typed events with quote and date; typed judgments for duplicate and conflict pairs and for pre-existing conditions; re-reads of regions the parser was unsure of","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"layout","role":"Finds and orders the regions of every page (text, headings, tables, checkboxes) with their boxes","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned page (text, handwriting, tables as cells)","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."},{"id":"retrieval","role":"Scans into retrieval: each page's elements as chunks with page and box, the index and layout hashes, every cite located to a chunk and byte span, and a reranked search per procedure for other pages that give another date","name":"Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)","where":"gpu","load":"resident","gb":9,"min_gb":9,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 9 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 34.4 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 42.4 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (72 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (72 of 192 GB)."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build."}}},{"id":"nsa-idr-packet","name":"No Surprises Act IDR packet and eligibility screen","page":"https://decosa.ai/clinics/nsa-idr-packet#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the claim facts from the EOB, remittance and notices (each with a quote), gathers the evidence for each factor the arbiter must consider, drafts the offer brief, and judges every brief sentence (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long files and dense correspondence","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long files and dense correspondence","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"walkthrough-to-quote","name":"Walkthrough-to-quote","page":"https://decosa.ai/tools/operations/walkthrough-to-quote#self-host","tiers":[{"id":"lite","label":"Lite · scope and quote from the video only","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.","name":"Qwen3.8-27B (NVIDIA NVFP4), video input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"quote","role":"Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)","name":"decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · narration, scope and quote (hosted demo)","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Watches the walkthrough (one video part per call, 1 frame a second; 2-minute parts past 150 s) and lists areas with dimension ranges and every visible condition with its time; reads the narration for requests and spoken measurements; drafts scope lines citing both; takes a second look at lines only the narration mentions. Never writes a quantity or a price.","name":"Qwen3.8-27B (NVIDIA NVFP4), video input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance; a long segment is split into sentences with approximate times","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"quote","role":"Clip preparation and chunking, keyframes, quantities (said numbers checked against the transcript, video ranges with their method, or missing), pricing from your sheet in integer cents with each unit price checked by the numeric-grounding block, measure-first ranking, re-pricing and the signed record (no model; CPU)","name":"decosa-api video block, numeric-grounding block and walkthrough module (decosa_api/video, decosa_api/verticals/numeric, decosa_api/verticals/walkthrough)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4), video input needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"expert-to-sop","name":"Expert-to-SOP","page":"https://decosa.ai/tools/operations/expert-to-sop#self-host","tiers":[{"id":"lite","label":"Lite · steps and keyframes, no narration check","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Watches the recording (one video part, sampled at 1 frame a second) and lists each step with its time range; judges each step against the narration near its time (the grounding block's judge); lists instructions said but not drafted, then checks them back against the video","name":"Qwen3.8-27B (NVIDIA NVFP4), video input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"sop","role":"Clip preparation (same timeline, no audio, at most 1280 px), keyframes at each cited time, statuses, sign-off and the signed revision record (no model; CPU)","name":"decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · steps checked against the narration (hosted demo)","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"llm","role":"Watches the recording (one video part, sampled at 1 frame a second) and lists each step with its time range; judges each step against the narration near its time (the grounding block's judge); lists instructions said but not drafted, then checks them back against the video","name":"Qwen3.8-27B (NVIDIA NVFP4), video input","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"The narration as timed lines, with a model-call receipt (audio hash in, transcript hash out) signed by the instance","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"sop","role":"Clip preparation (same timeline, no audio, at most 1280 px), keyframes at each cited time, statuses, sign-off and the signed revision record (no model; CPU)","name":"decosa-api video block and SOP module (decosa_api/video, decosa_api/verticals/sop)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4), video input needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4), video input: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4), video input with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"honest-product-imagery","name":"Honest product imagery","page":"https://decosa.ai/tools/media/honest-product-imagery#self-host","tiers":[{"id":"lite","label":"Lite · pixels and the model, no document reader","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Finds the product units, the logo and label, and every person or human likeness in the reference photo and the AI image; compares the two side by side with the seller's facts, per category (size, colour, text, logo, warning, feature)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"check","role":"Alignment (NCC on edges over scale and a few degrees of rotation), colour (CIEDE2000 after white balance on the pack's neutral areas), logo match, unit count, changed pack area, the quantity and text comparisons, the verdict rules, the consent gate, XMP, the burned-in label and its OCR read-back, the C2PA credential and the signed record (no model; CPU)","name":"decosa-api imagery module (decosa_api/verticals/imagery)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · pixels, pack text and the model (hosted demo)","gpu_gb":63,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Finds the product units, the logo and label, and every person or human likeness in the reference photo and the AI image; compares the two side by side with the seller's facts, per category (size, colour, text, logo, warning, feature)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"layout","role":"Pack text, step 1: finds the text regions on crops of the product (the document reader block)","name":"Docling 2.130 with the Heron layout model","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Pack text, step 2: reads each region (name, variant, claims, quantity, warnings)","name":"PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."},{"id":"check","role":"Alignment (NCC on edges over scale and a few degrees of rotation), colour (CIEDE2000 after white balance on the pack's neutral areas), logo match, unit count, changed pack area, the quantity and text comparisons, the verdict rules, the consent gate, XMP, the burned-in label and its OCR read-back, the C2PA credential and the signed record (no model; CPU)","name":"decosa-api imagery module (decosa_api/verticals/imagery)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"alternate-demo-scenes","label":"Lifestyle scenes for demos (not part of the service)","gpu_gb":34,"basis":"stack","unknown":[],"components":[{"id":"render","role":"Not part of the service: drew the demo and eval lifestyle scenes around synthetic packs (one job at a time on the shared studio GPU)","name":"Wan2.2-VACE-Fun-A14B","where":"gpu","load":"job","gb":34,"min_gb":34,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"Wan 2.2 VACE A14B (FP8): About 34 GB of GPU for the renderer while a job runs (animatic-studio stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 25.4 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Docling 2.130 with the Heron layout model has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Docling 2.130 with the Heron layout model has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"pv-intake","name":"Pharmacovigilance intake","page":"https://decosa.ai/tools/life-sciences/pv-intake#self-host","tiers":[{"id":"lite","label":"Lite · one 48-80 GB card","gpu_gb":89.5,"basis":"estimate","unknown":[],"components":[{"id":"pv","role":"Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)","name":"decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm-lite","role":"The same prompts on a smaller mixture-of-experts model","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":57,"min_gb":53,"weights_gb":49,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, BF16: BF16 weights are 49 GB (stack.json); the KV cache on top is an estimate."},{"id":"asr","role":"Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate. (stack.json lists 3 GB for this component.)"},{"id":"asr-ml","role":"Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)","name":"Qwen3-ASR-1.7B (language pack)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"mt","role":"Reports not in English, translated line by line to English, and follow-up questions back into the reporter's language","name":"Hy-MT2-7B (language pack)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"layout","role":"Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positions","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned form","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]},{"id":"standard","label":"Standard · the hosted demo","gpu_gb":90.1,"basis":"estimate","unknown":[],"components":[{"id":"pv","role":"Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)","name":"decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Reads the fields with quotes, judges the seriousness criteria per event and expectedness against the label, drafts the narrative and judges its sentences (grounding)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"asr","role":"Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate. (stack.json lists 3 GB for this component.)"},{"id":"asr-ml","role":"Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)","name":"Qwen3-ASR-1.7B (language pack)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"mt","role":"Reports not in English, translated line by line to English, and follow-up questions back into the reporter's language","name":"Hy-MT2-7B (language pack)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"layout","role":"Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positions","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned form","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]},{"id":"best","label":"Best · a larger judge","gpu_gb":224.5,"basis":"estimate","unknown":[],"components":[{"id":"pv","role":"Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)","name":"decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm-best","role":"A larger judge for seriousness and expectedness","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate. (stack.json lists 3 GB for this component.)"},{"id":"asr-ml","role":"Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)","name":"Qwen3-ASR-1.7B (language pack)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"mt","role":"Reports not in English, translated line by line to English, and follow-up questions back into the reporter's language","name":"Hy-MT2-7B (language pack)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"layout","role":"Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positions","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned form","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]},{"id":"wanted","label":"Wanted: the best setup · a large judge with room to spare","gpu_gb":224.5,"basis":"estimate","unknown":[],"components":[{"id":"pv","role":"Intake, the quote checks, the four minimum criteria, day 0 and the clocks, follow-up questions, the E2B(R3)-shaped draft, signing and the HTTP API (/pv/*)","name":"decosa-api pharmacovigilance intake (decosa_api/verticals/pv), importing the language pack, the document reader, grounding (17), typed judgment (24) and the dates block (54)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm-best","role":"A larger judge for seriousness and expectedness","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"asr","role":"Call recordings in English to a timed, speaker-labelled transcript (the fallback for other languages)","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate. (stack.json lists 3 GB for this component.)"},{"id":"asr-ml","role":"Call recordings in German, French, Spanish, Italian, Dutch, Polish, Portuguese or Czech (the fallback for English)","name":"Qwen3-ASR-1.7B (language pack)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"mt","role":"Reports not in English, translated line by line to English, and follow-up questions back into the reporter's language","name":"Hy-MT2-7B (language pack)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"layout","role":"Finds and orders the regions of a scanned form (text, boxes, checkboxes) with their positions","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads each region of a scanned form","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 52.5 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 60.5 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 66.1 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (90.1 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (90.1 of 192 GB)."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack) has no mapped Apple Silicon build; Hy-MT2-7B (language pack) has no mapped Apple Silicon build; Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack) has no mapped Apple Silicon build; Hy-MT2-7B (language pack) has no mapped Apple Silicon build; Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build."}}},{"id":"label-consistency-check","name":"Label consistency across PI, SmPC, CCDS and carton","page":"https://decosa.ai/tools/life-sciences/label-consistency-check#self-host","tiers":[{"id":"lite","label":"Lite · text documents, meaning check off","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"labelcheck","role":"Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)","name":"decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)","name":"Qwen3.8-27B (NVFP4, vision tower on)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · Qwen3.8-27B, Hy-MT2-7B and the document reader (hosted demo)","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"labelcheck","role":"Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)","name":"decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)","name":"Qwen3.8-27B (NVFP4, vision tower on)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"mt","role":"Back-translation of each translated paragraph into English for the meaning check (the language-pack block): German, French and the other languages it serves","name":"Hy-MT2-7B","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"reader","role":"Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per line","name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."}]},{"id":"best","label":"Best · the 30B translation model (not measured)","gpu_gb":100.6,"basis":"estimate","unknown":[],"components":[{"id":"labelcheck","role":"Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)","name":"decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)","name":"Qwen3.8-27B (NVFP4, vision tower on)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"mt-30b","role":"Back-translation, larger model","name":"Hy-MT2-30B-A3B-FP8","where":"gpu","load":"resident","gb":37,"min_gb":37,"weights_gb":30,"basis":"estimate","precision":"fp8","gpus":1,"source":"Estimate: 30B parameters at 1 byte per weight is about 30 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"reader","role":"Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per line","name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4, vision tower on) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 44 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4, vision tower on): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 57.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4, vision tower on) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (81.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (81.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Hy-MT2-7B has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Hy-MT2-7B has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"no-training-receipts","name":"No-training receipts","page":"https://decosa.ai/data","tiers":[{"id":"lite","label":"Lite · check a report on your own machine","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"verifier","role":"The offline verifier: rechecks signatures, receipts, counts, retention, the ledger chain and the cross-check; with --my-data it hashes your own documents on your machine and looks them up in every dataset","name":"verify.py (one file; Python standard library plus cryptography)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the report in decosa-api (CPU)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"report","role":"Builds the signed report: collects the key's receipts for the period, the retention table and the handling events, cross-checks the calls against the training ledger, and signs (no model; CPU)","name":"decosa-api no-training receipts (decosa_api/verticals/notrain)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ledger","role":"The training-job ledger: every training run is recorded before it reads data (dataset hashes, declared sources and licences, fingerprints) and when it ends (output weights hashes)","name":"Signed training-job ledger (decosa.training-run.v1) with scripts/notrain_train.py and the QE trainer's hook","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"verifier","role":"The offline verifier: rechecks signatures, receipts, counts, retention, the ledger chain and the cross-check; with --my-data it hashes your own documents on your machine and looks them up in every dataset","name":"verify.py (one file; Python standard library plus cryptography)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · the same report over confidential-tier calls","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"report","role":"Builds the signed report: collects the key's receipts for the period, the retention table and the handling events, cross-checks the calls against the training ledger, and signs (no model; CPU)","name":"decosa-api no-training receipts (decosa_api/verticals/notrain)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"ledger","role":"The training-job ledger: every training run is recorded before it reads data (dataset hashes, declared sources and licences, fingerprints) and when it ends (output weights hashes)","name":"Signed training-job ledger (decosa.training-run.v1) with scripts/notrain_train.py and the QE trainer's hook","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"verifier","role":"The offline verifier: rechecks signatures, receipts, counts, retention, the ledger chain and the cross-check; with --my-data it hashes your own documents on your machine and looks them up in every dataset","name":"verify.py (one file; Python standard library plus cryptography)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 0 GB). The best tier fits too."},"rtx-4090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 24 GB). The best tier fits too."},"rtx-5090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 32 GB). The best tier fits too."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 64 GB). The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 80 GB). The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 72 GB). The best tier fits too."},"m5-max-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."}}},{"id":"storefront-accessibility-pass","name":"Storefront accessibility pass","page":"https://decosa.ai/tools/operations/storefront-accessibility-pass#self-host","tiers":[{"id":"lite","label":"Lite · rules, keyboard and forms, no model (CPU)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"browser","role":"Opens each page in a private headless Chromium (one per audit), runs axe-core, the keyboard pass (Tab order, Enter on add-to-cart, visible focus, Escape from dialogs) and the empty-submit form pass; blocks every request that is not GET or HEAD","name":"axe-core 4.13.0 in headless Chromium (Playwright 1.58)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"report","role":"Merges the rule, keyboard, form and model results into findings ranked P1 (blocks a purchase) to P4, cites the WCAG 2.2 success criteria, writes a code fix per finding and seals the signed record","name":"decosa-api a11y module (decosa_api/verticals/a11y)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · rules, keyboard, forms and the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Looks at each image with its alt text (right, wrong, poor, needed but empty, decorative), reads offer text drawn into banners, judges link and button names in their context and the form's error messages","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"browser","role":"Opens each page in a private headless Chromium (one per audit), runs axe-core, the keyboard pass (Tab order, Enter on add-to-cart, visible focus, Escape from dialogs) and the empty-submit form pass; blocks every request that is not GET or HEAD","name":"axe-core 4.13.0 in headless Chromium (Playwright 1.58)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"report","role":"Merges the rule, keyboard, form and model results into findings ranked P1 (blocks a purchase) to P4, cites the WCAG 2.2 success criteria, writes a code fix per finding and seals the signed record","name":"decosa-api a11y module (decosa_api/verticals/a11y)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVIDIA NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"fill-and-stop","name":"Fill a form from your papers","page":"https://decosa.ai/tools/operations/fill-and-stop#self-host","tiers":[{"id":"lite","label":"Lite · text and PDF sources, English review","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Maps each form field to the line of your documents that answers it (12 fields per call), checks page text the patterns did not settle for instructions aimed at AI agents, and reads a widget from a screenshot when the page's code cannot set it","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"browser","role":"A throwaway headless Chromium per run (no cookies, no storage, no service workers, no downloads): reads the form, types and chooses in the page, holds every request that could send the form and blocks every other host","name":"Playwright 1.58 with Chromium (headless)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"commit-detector","role":"Says whether a click the model chooses would commit something (pay, delete, send, publish, submit, security) and labels the buttons left for you; the first line, with the word list and the network hold always on","name":"decosa-commit-detector (XLM-RoBERTa-large fine-tune, our own model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"fillstop","role":"Field reader, the value checks (the value must be in the cited line, or the same date, time, phone number or amount written another way), the sensitive-field rules, the commit words in 32 languages, the injection patterns, the review, the signed approval and the release of the held Submit","name":"decosa-api fill-and-stop module (decosa_api/verticals/fillstop) on the shared computer-use engine (decosa_api.cu) and the flight recorder (use case 26)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · adds scans, photos and a review in 24 languages (hosted demo)","gpu_gb":81.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Maps each form field to the line of your documents that answers it (12 fields per call), checks page text the patterns did not settle for instructions aimed at AI agents, and reads a widget from a screenshot when the page's code cannot set it","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"reader","role":"Reads scanned PDFs and photos of papers (a police report photographed on a phone) into numbered lines the fields can point at","name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."},{"id":"mt","role":"Translates the review (field labels, reasons and your cited document lines) into the person's language; the values typed into the form are never translated","name":"Hy-MT2-7B (language pack; Qwen3.8-27B for the EU languages Hy-MT2 does not cover)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"browser","role":"A throwaway headless Chromium per run (no cookies, no storage, no service workers, no downloads): reads the form, types and chooses in the page, holds every request that could send the form and blocks every other host","name":"Playwright 1.58 with Chromium (headless)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"commit-detector","role":"Says whether a click the model chooses would commit something (pay, delete, send, publish, submit, security) and labels the buttons left for you; the first line, with the word list and the network hold always on","name":"decosa-commit-detector (XLM-RoBERTa-large fine-tune, our own model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"fillstop","role":"Field reader, the value checks (the value must be in the cited line, or the same date, time, phone number or amount written another way), the sensitive-field rules, the commit words in 32 languages, the injection patterns, the review, the signed approval and the release of the held Submit","name":"decosa-api fill-and-stop module (decosa_api/verticals/fillstop) on the shared computer-use engine (decosa_api.cu) and the flight recorder (use case 26)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"wanted","label":"Wanted: the best setup · a second model on choices and widgets","gpu_gb":182.9,"basis":"estimate","unknown":[],"components":[{"id":"llm-bf16-x2","role":"A larger vision-language model as a second opinion on the choices marked 'check it' and on widgets (not built yet)","name":"Qwen3.8-27B in BF16 alongside a grounding model (Holo3-35B-A3B, Apache-2.0) for widgets","where":"gpu","load":"resident","gb":158.9,"min_gb":158.9,"weights_gb":131.6,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 65.8B parameters at 2 bytes (BF16) per weight is about 131.6 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"reader","role":"Reads scanned PDFs and photos of papers (a police report photographed on a phone) into numbered lines the fields can point at","name":"Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."},{"id":"mt","role":"Translates the review (field labels, reasons and your cited document lines) into the person's language; the values typed into the form are never translated","name":"Hy-MT2-7B (language pack; Qwen3.8-27B for the EU languages Hy-MT2 does not cover)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"browser","role":"A throwaway headless Chromium per run (no cookies, no storage, no service workers, no downloads): reads the form, types and chooses in the page, holds every request that could send the form and blocks every other host","name":"Playwright 1.58 with Chromium (headless)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"commit-detector","role":"Says whether a click the model chooses would commit something (pay, delete, send, publish, submit, security) and labels the buttons left for you; the first line, with the word list and the network hold always on","name":"decosa-commit-detector (XLM-RoBERTa-large fine-tune, our own model)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"fillstop","role":"Field reader, the value checks (the value must be in the cited line, or the same date, time, phone number or amount written another way), the sensitive-field rules, the commit words in 32 languages, the injection patterns, the review, the signed approval and the release of the held Submit","name":"decosa-api fill-and-stop module (decosa_api/verticals/fillstop) on the shared computer-use engine (decosa_api.cu) and the flight recorder (use case 26)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 44 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 57.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (81.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B) has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B) has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"privileged-call-notes","name":"Privileged call notes","page":"https://decosa.ai/legal/privileged-call-notes#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card, uploaded calls","gpu_gb":37.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-pass2","role":"After hang-up (or on an upload): who said what, one line per turn, each with its own receipt","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm-lite","role":"Lite tier: the same text steps on the official FP8 checkpoint","name":"Qwen3.8-27B (official FP8)","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers)."}]},{"id":"standard","label":"Standard · the hosted demo","gpu_gb":85.6,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions during the call (streaming, no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After hang-up (or on an upload): who said what, one line per turn, each with its own receipt","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm","role":"Speaker roles, live intake checklist, cited memo, conflicts names, claim check, time-entry narrative and client email","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash writes and checks","gpu_gb":220,"basis":"estimate","unknown":[],"components":[{"id":"asr-live","role":"Live captions during the call (streaming, no speakers)","name":"Voxtral Mini 4B Realtime","where":"gpu","load":"resident","gb":24,"min_gb":16,"weights_gb":8.3,"basis":"stack","precision":null,"gpus":1,"source":"Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more."},{"id":"asr-pass2","role":"After hang-up (or on an upload): who said what, one line per turn, each with its own receipt","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"llm-best","role":"Best tier: memo writer and claim checker","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Voxtral Mini 4B Realtime needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 40 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 48 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (85.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Voxtral Mini 4B Realtime with Voxtral Mini 4B Realtime, MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 48 GB available to the GPU. The lite tier fits with changes."}}},{"id":"check-their-brief","name":"Check their brief","page":"https://decosa.ai/legal/check-their-brief#self-host","tiers":[{"id":"lite","label":"Lite · CPU only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"checker","role":"Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)","name":"decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · one GPU for the judge (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":["reader"],"components":[{"id":"checker","role":"Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)","name":"decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"reader","role":"Document reader for scanned filings (no text layer): layout plus OCR, then the same checks","name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","where":"gpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Memory not stated in stack.json and not derivable (no parameter count)."}]},{"id":"wanted","label":"Wanted · a GLM-5.3-Flash holdings judge on your own hardware","gpu_gb":249.6,"basis":"estimate","unknown":[],"components":[{"id":"checker","role":"Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)","name":"decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"glm-judge","role":"Stronger holdings judge (wanted)","name":"GLM-5.3-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]},{"id":"alternate-support-model","label":"Our citation-support model as a second opinion on holdings","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"support","role":"Citation-support model (own model M1, prototype): does the opinion's passage support the brief's sentence","name":"decosa-citation-support-modernbert-large (own model M1, prototype; Apache-2.0)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[{"id":"reader","name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","tiers":["standard"]}],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"l40sx1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"h100x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"rtx-pro-6000x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"rtx-pro-6000x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits."}}},{"id":"discovery-deficiency","name":"Check their discovery responses","page":"https://decosa.ai/legal/discovery-deficiency#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"discovery","role":"Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)","name":"decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shape","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · Qwen3.8-27B and the document reader (hosted demo)","gpu_gb":63.6,"basis":"stack","unknown":[],"components":[{"id":"discovery","role":"Splitter, set checks (verification, signature, deadlines), flag rules, quote location, rules pack, letter and fix list, signed record (no model; CPU)","name":"decosa-api discovery check (decosa_api/verticals/discovery), with the dates block, the drafting editor's Word writer and the signed record (07)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Reads each response once and describes it as JSON: objection grounds and whether each gives specifics, withholding statement, production date, answer shape, admission shape","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"reader","role":"Scanned PDFs only: page images to text","name":"Document reader (Docling layout heron + PaddleOCR-VL-1.6)","where":"gpu","load":"resident","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 6 in stack.json."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 26 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 34 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Document reader (Docling layout heron + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Document reader (Docling layout heron + PaddleOCR-VL-1.6) has no mapped Apple Silicon build The lite tier fits with changes."}}},{"id":"payer-audit","name":"Payer audit response","page":"https://decosa.ai/clinics/payer-audit#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the auditor's letter (who, dates, reference, policy named), reads a pasted policy's requirements when no pack is given, and for each claim points at the note's times, signature and addenda and judges each content requirement found or missing with the exact words; drafts the cover letter body and judges each of its sentences (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"prior-auth-check","name":"Prior-auth pre-check and packet","page":"https://decosa.ai/clinics/prior-auth-check#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same pipeline on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Picks the criteria section of a long policy, splits the policy into requirements, alternatives and exclusions, checks each against the chart, re-checks every not-met answer (and every exclusion answered met), answers the payer's form questions, drafts the letter of medical necessity when the chart supports every criterion, and judges every letter sentence (the grounding judge)","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long charts and dense payer policies","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long charts and dense payer policies","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"jottings-note","name":"Notes from your own jottings","page":"https://decosa.ai/clinics/jottings-note#self-host","tiers":[{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"note","role":"Reading typed jottings, the completeness check (dates, times and risk in code), the number, clinical-claim, qualifier and attribution guards, layouts, catch-up and the signed record (no model; CPU)","name":"decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Drafts the sentences, tags the audited elements, checks every sentence (the grounding judge), rewrites a failed sentence once, and reads a photo's page","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"best","label":"Best · adds the second reader and dictation","gpu_gb":65.7,"basis":"estimate","unknown":[],"components":[{"id":"note","role":"Reading typed jottings, the completeness check (dates, times and risk in code), the number, clinical-claim, qualifier and attribution guards, layouts, catch-up and the signed record (no model; CPU)","name":"decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Drafts the sentences, tags the audited elements, checks every sentence (the grounding judge), rewrites a failed sentence once, and reads a photo's page","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"reader","role":"Second reader for handwriting: reads each line of a photo again, so lines the two readers read differently are flagged for the clinician","name":"PaddleOCR-VL-1.6 (the document reader's page parser)","where":"gpu","load":"resident","gb":3,"min_gb":3,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 3 in stack.json."},{"id":"asr","role":"Transcribes a dictation of up to 60 seconds made after the session (never a session recording)","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size). The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"evidence-runner","name":"Capture audit evidence from your admin screens","page":"https://decosa.ai/tools/finance/evidence-runner#self-host","tiers":[{"id":"lite","label":"Lite · quarterly runs only, no GPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"runner","role":"Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)","name":"decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · setup and repairs with Qwen3.8-27B","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)","name":"decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"Setup and repairs only: finds each named screen (read-only) and points at rows when code cannot place a setting","name":"Qwen3.8-27B","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 96 GB for this component.)"}]},{"id":"alternate-api-export","label":"Official read-only API exports (Graph, Policy API, AWS, Okta)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"runner","role":"Browser session with the read-only gate, saved routes, the code reader for named settings, stamped captures, certificate and evidence pack (CPU)","name":"decosa-api evidence runner (decosa_api/verticals/evidence) on the computer-use engine (decosa_api.cu) and the test-run certificate (27)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Qwen3.8-27B needs a GPU. The lite tier fits."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B: run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"ehr-drafts","name":"Put the visit into your EHR as drafts","page":"https://decosa.ai/clinics/ehr-drafts#self-host","tiers":[{"id":"standard","label":"Standard · one GPU for the model","gpu_gb":33.6,"basis":"stack","unknown":[],"components":[{"id":"agent","role":"The agent that fills the EHR's forms: one receipted decision per step, values only from the visit, the one draft save released after the chart and form are checked in code","name":"Qwen3.8-27B on the Decosa computer-use engine","where":"gpu","load":"resident","gb":33.6,"min_gb":32,"weights_gb":29,"basis":"stack","precision":"fp8","gpus":1,"source":"Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers). (stack.json lists 20 GB for this component.)"},{"id":"medcheck","role":"Reads each medicine (drug, dose, frequency, route when said, duration) against the transcript lines it came from","name":"decosa-note-detail-checker (M17)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"gate","role":"Holds sign, finalise, transmit, send, fax, e-prescribe, bill and delete requests in the browser; releases each draft save once","name":"decosa-api ehr-drafts and the computer-use network gate","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B on the Decosa computer-use engine needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). Qwen3.8-27B on the Decosa computer-use engine needs about 32 GB at its smallest setting. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B on the Decosa computer-use engine: run it at its smallest setting (about 32 GB instead of 33.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (33.6 of 48 GB)."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (33.6 of 80 GB)."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (33.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (33.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B on the Decosa computer-use engine with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"demand-reader","name":"Injury demand reader","page":"https://decosa.ai/tools/insurance/demand-reader#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same reads on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the letter's terms and every condition with quotes, each bill page's charge lines, and each record page's visits (the medical chronology's page extractor); the grounding judge checks every sentence it writes","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for packages of several hundred pages","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for packages of several hundred pages","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"certificate-check","name":"Certificate request check","page":"https://decosa.ai/tools/insurance/certificate-check#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same extraction and checks on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Lists each requirement from the contract and the email with a verbatim quote, then per requirement says whether the policy papers meet it, quoting the policy; the grounding judge checks each reason","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long contracts with many exhibits","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long contracts with many exhibits","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"bank-change-check","name":"Vendor bank-change check","page":"https://decosa.ai/tools/finance/bank-change-check#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same quoted reading on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Reads the email's text and quotes the pressure, secrecy, 'don't call' and redirected-payment signs, and reads the new account's bank and holder; the header, domain and vendor-file checks are code","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, forwarded email threads","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Best tier: a larger model for long, forwarded email threads","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"family-film","name":"Family interview film","page":"https://decosa.ai/apps/family-film#self-host","tiers":[{"id":"lite","label":"Lite · chapters, quotes, subtitles and the film, no photo reading","gpu_gb":80.7,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Speech recognition in her language, one call per speech segment (so every word keeps its time)","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"speakers","role":"Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hers","name":"ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallback","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"translate","role":"English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every line","name":"Hy-MT2-7B (the language-pack block)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"music","role":"A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)","name":"ACE-Step 1.5 cue (music-gen-cleared library)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"edit","role":"The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per file","name":"decosa-api family_film module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, with photo backs read","gpu_gb":80.7,"basis":"estimate","unknown":["reader"],"components":[{"id":"asr","role":"Speech recognition in her language, one call per speech segment (so every word keeps its time)","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"speakers","role":"Who is speaking: voice activity by energy, then ECAPA-TDNN voice embeddings per segment in two clusters; the cluster nearest her consent recording is hers","name":"ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb, ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Chapters of her life and her best lines (copied exactly from the transcript, then re-found in it by code), film titles, and the translation fallback","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"translate","role":"English subtitles under her own words, sentence by sentence, then a meaning check (back-translation compared with the source) on every line","name":"Hy-MT2-7B (the language-pack block)","where":"gpu","load":"resident","gb":18,"min_gb":18,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 18 in stack.json."},{"id":"faces","role":"Face detection only (boxes): a picture is read as the back of a photo only when no face is found; photo pans drift toward faces","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"reader","role":"Reads the handwriting on the backs of photos (place, year, names) to date and place each picture","name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","where":"gpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Memory not stated in stack.json and not derivable (no parameter count)."},{"id":"music","role":"A quiet score under the film: a pre-rendered cue from the cleared music library (no model runs per film)","name":"ACE-Step 1.5 cue (music-gen-cleared library)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"edit","role":"The film (CPU): her voice over her photos with slow pans, chapter cards, maps (Natural Earth) and dates, subtitles, the trailer and the book with a QR code per quote; C2PA credential per file","name":"decosa-api family_film module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[{"id":"reader","name":"Decosa document reader (Docling layout + PaddleOCR-VL-1.6)","tiers":["standard"]}],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3-ASR-1.7B (language pack speech service) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 43.1 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 51.1 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits with changes."},"l40sx1":{"verdict":"no","recommended":null,"reason":"Needs about 56.7 GB of GPU memory at the smallest settings; 48 GB available."},"h100x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits with changes."},"rtx-pro-6000x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"rtx-pro-6000x2":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (memory not known) The lite tier fits."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; Hy-MT2-7B (the language-pack block) has no mapped Apple Silicon build; Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; Hy-MT2-7B (the language-pack block) has no mapped Apple Silicon build; Decosa document reader (Docling layout + PaddleOCR-VL-1.6) has no mapped Apple Silicon build."}}},{"id":"music-video-starring-you","name":"Music video starring you","page":"https://decosa.ai/apps/music-video-starring-you#self-host","tiers":[{"id":"standard","label":"Standard · the hosted demo, a storyboard cut on the beat","gpu_gb":68.7,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"faces","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checks","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"analyzer","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"stills","role":"Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)","name":"FLUX.2 klein 4B","where":"gpu","load":"job","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":"fp4","gpus":1,"source":"vram_gb 6 in stack.json."},{"id":"edit","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","name":"decosa-api starring module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · moving shots on MiniMax H3, one 96 GB card","gpu_gb":52,"basis":"measured","unknown":[],"components":[{"id":"asr","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"faces","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Adult check on the consent frames (vision), the shot list per song section, and the frame safety and likeness checks","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"analyzer","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"h3","role":"Moving shots: MiniMax H3 reference mode with the performer's consent-clip frames, Turbo v4 LoRA, sketch tier 864x480","name":"MiniMax-H3 (reference mode, Turbo v4)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":"fp8","gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json). (stack.json lists 52 GB for this component.)"},{"id":"edit","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","name":"decosa-api starring module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3-ASR-1.7B (language pack speech service) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"h100x1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (68.7 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (68.7 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build."}}},{"id":"our-story-film","name":"Our story film","page":"https://decosa.ai/apps/our-story-film#self-host","tiers":[{"id":"standard","label":"Standard · the hosted demo, a storyboard cut on the beat","gpu_gb":68.7,"basis":"estimate","unknown":[],"components":[{"id":"asr","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"faces","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Adult check on the consent frames (vision), the shot list per song section from your three memories, and the frame safety and likeness checks","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"analyzer","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"stills","role":"Storyboard: one still per shot drawn from the consent-clip references, then a camera move (push, pan, drift)","name":"FLUX.2 klein 4B","where":"gpu","load":"job","gb":6,"min_gb":6,"weights_gb":null,"basis":"stack","precision":"fp4","gpus":1,"source":"vram_gb 6 in stack.json."},{"id":"edit","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","name":"decosa-api starring module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · moving shots on MiniMax H3, one 96 GB card","gpu_gb":52,"basis":"measured","unknown":[],"components":[{"id":"asr","role":"Consent read-back: hears whether the clip says the sentence and its three fresh words","name":"Qwen3-ASR-1.7B (language pack speech service)","where":"gpu","load":"resident","gb":5.1,"min_gb":5.1,"weights_gb":3.4,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 1.7B parameters at 2 bytes (BF16) per weight is about 3.4 GB, plus 20% working memory and 1 GB of runtime. Not measured."},{"id":"faces","role":"Face detection only (boxes): exactly one face, present and moving through the clip; its three sharpest frames become the only face references","name":"Ultra-Light-Fast-Generic-Face-Detector-1MB (version-RFB-320)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Adult check on the consent frames (vision), the shot list per song section from your three memories, and the frame safety and likeness checks","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"analyzer","role":"Beat grid, bars and sections of the song (CPU; the music-video studio's analyzer): cuts land on bar downbeats","name":"decosa-mvideo-analyze (services/mvideo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"h3","role":"Moving shots: MiniMax H3 reference mode with the performer's consent-clip frames (two subjects), Turbo v4 LoRA, sketch tier 864x480","name":"MiniMax-H3 (reference mode, Turbo v4)","where":"gpu","load":"job","gb":96,"min_gb":96,"weights_gb":null,"basis":"stack","precision":"fp8","gpus":1,"source":"MiniMax-H3 (self-host): A whole 96 GB card plus about 115 GB of system RAM for CPU offload (studio and ugc stack.json). (stack.json lists 52 GB for this component.)"},{"id":"edit","role":"The edit (CPU): shots cut on the bar downbeats at 30 fps, the AI video label on every frame, an end card crediting the music, exports in 16:9, 9:16 and a Spotify Canvas loop; cut timing measured back from the pixels; C2PA per file","name":"decosa-api starring module + FFmpeg + c2pa-python","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3-ASR-1.7B (language pack speech service) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 39.1 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"h100x1":{"verdict":"no","recommended":null,"reason":"FLUX.2 klein 4B: This build is NVIDIA NVFP4, which needs a Blackwell GPU. No replacement is listed."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (68.7 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (68.7 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Qwen3-ASR-1.7B (language pack speech service) has no mapped Apple Silicon build; FLUX.2 klein 4B has no mapped Apple Silicon build."}}},{"id":"settlement-video","name":"Settlement video from the case file","page":"https://decosa.ai/legal/settlement-video#self-host","tiers":[{"id":"lite","label":"Lite · captions or your own recording, no voice models","gpu_gb":63,"basis":"stack","unknown":[],"components":[{"id":"settlement","role":"The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)","name":"decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure of","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"layout","role":"Finds the regions of each bill and statement page (tables, text) with their boxes","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads the bill tables as cells and the statement's paragraphs","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."}]},{"id":"standard","label":"Standard · reader, model, house voices and the consented clone (hosted demo)","gpu_gb":63,"basis":"stack","unknown":[],"components":[{"id":"settlement","role":"The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)","name":"decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocks","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"model","role":"The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure of","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"layout","role":"Finds the regions of each bill and statement page (tables, text) with their boxes","name":"Docling 2.130 with the Heron layout model (document reader block)","where":"gpu","load":"resident","gb":1,"min_gb":1,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 1 in stack.json."},{"id":"parser","role":"Reads the bill tables as cells and the statement's paragraphs","name":"PaddleOCR-VL-1.6 (0.9B, document reader block)","where":"gpu","load":"resident","gb":4.4,"min_gb":4.4,"weights_gb":null,"basis":"stack","precision":null,"gpus":1,"source":"vram_gb 4.4 in stack.json."},{"id":"house-voice","role":"The stock house voice that reads the approved script, when the lawyer picks it","name":"Kokoro-82M (stock voicepacks am_michael, af_heart)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"consented-voice","role":"The attorney's own voice, generated from the attorney's consent recording, only with an active consent-ledger entry for this matter","name":"Chatterbox Multilingual","where":"cpu","load":"job","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"no","recommended":null,"reason":"Needs about 25.4 GB of GPU memory at the smallest settings; 24 GB available."},"rtx-5090x1":{"verdict":"no","recommended":null,"reason":"Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (63 of 192 GB)."},"m3-ultra-96x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build."},"m5-max-64x1":{"verdict":"unknown","recommended":null,"reason":"Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build."}}},{"id":"interview-themes","name":"Interview themes","page":"https://decosa.ai/tools/research/interview-themes#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":44,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Lite tier: the same codebook, coding and themes on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."},{"id":"diarizer","role":"Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute piece","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"voice","role":"Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"embedder","role":"Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)","name":"bge-small-en-v1.5 (ONNX)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":61.6,"basis":"estimate","unknown":[],"components":[{"id":"diarizer","role":"Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute piece","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"voice","role":"Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Proposes the codebook, applies the approved codebook to every passage of participant talk (with the words that justify each code), groups codes into themes and picks candidate quotes; participant counts and the quote check are code","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"},{"id":"embedder","role":"Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)","name":"bge-small-en-v1.5 (ONNX)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"wanted","label":"Wanted · a much larger coder on your own hardware","gpu_gb":196,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Wanted: a much larger coder for long studies and finer codebooks","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"diarizer","role":"Recordings in: speech recognition with speaker turns, one pass per 6 to 10 minute piece","name":"MOSS-Transcribe-Diarize 0.9B","where":"gpu","load":"resident","gb":4,"min_gb":4,"weights_gb":1.8,"basis":"estimate","precision":null,"gpus":1,"source":"MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate."},{"id":"voice","role":"Voice check: links each piece's speakers into one voice per person for the whole recording, and moves segments whose voice matches the other speaker (marked in the transcript)","name":"ECAPA-TDNN speaker embeddings (ONNX export)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"embedder","role":"Your own model: sentence embeddings for the small classifier trained on your reviewed codes (one logistic-regression head per code)","name":"bge-small-en-v1.5 (ONNX)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"MOSS-Transcribe-Diarize 0.9B needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (61.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace MOSS-Transcribe-Diarize 0.9B with MOSS-Transcribe-Diarize, MLX 8-bit. MLX build for Apple Silicon. (Memory is an estimate.)"},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace MOSS-Transcribe-Diarize 0.9B with MOSS-Transcribe-Diarize, MLX 8-bit. MLX build for Apple Silicon. (Memory is an estimate.)"}}},{"id":"mix-cue-sheet","name":"Mix cue sheet","page":"https://decosa.ai/tools/media/mix-cue-sheet#self-host","tiers":[{"id":"lite","label":"Lite · titles and times only, any CPU","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"cue-engine","role":"Song starts: spectral features and a novelty curve from the audio, a tracklist parser, and a resolver that turns your times, your track-file matches and the audio's change points into one start per song; writes the chapters (no model; CPU)","name":"decosa-cue engine (decosa_api/cue)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"standard","label":"Standard · adds your own track files (hosted)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"cue-engine","role":"Song starts: spectral features and a novelty curve from the audio, a tracklist parser, and a resolver that turns your times, your track-file matches and the audio's change points into one start per song; writes the chapters (no model; CPU)","name":"decosa-cue engine (decosa_api/cue)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"landmarks","role":"Finds each of your own track files in the mix from the audio (landmark pairs, searched across tempo changes of up to 6% and the pitch shift that comes with them) so its song gets an exact start (no model; CPU)","name":"Landmark matcher for your own track files (decosa_api/verticals/clearance fingerprinter)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]},{"id":"best","label":"Best · adds the learned transition detector (self-host only)","gpu_gb":0,"basis":null,"unknown":[],"components":[{"id":"cue-engine","role":"Song starts: spectral features and a novelty curve from the audio, a tracklist parser, and a resolver that turns your times, your track-file matches and the audio's change points into one start per song; writes the chapters (no model; CPU)","name":"decosa-cue engine (decosa_api/cue)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"landmarks","role":"Finds each of your own track files in the mix from the audio (landmark pairs, searched across tempo changes of up to 6% and the pitch shift that comes with them) so its song gets an exact start (no model; CPU)","name":"Landmark matcher for your own track files (decosa_api/verticals/clearance fingerprinter)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"detector","role":"Transition detector v1: a small dilated 1-D CNN that scores where one track hands over to the next, used in place of the novelty curve (self-host only, off by default)","name":"Transition detector v1 (ONNX)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 0 GB). The best tier fits too."},"rtx-4090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 24 GB). The best tier fits too."},"rtx-5090x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 32 GB). The best tier fits too."},"rtx-5090x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 64 GB). The best tier fits too."},"l40sx1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."},"h100x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 80 GB). The best tier fits too."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 96 GB). The best tier fits too."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 72 GB). The best tier fits too."},"m5-max-64x1":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (0 of 48 GB). The best tier fits too."}}},{"id":"geo-audit","name":"Site readiness for AI answers","page":"https://decosa.ai/family#elmoseo","tiers":[{"id":"lite","label":"Lite · no local GPU, the judge runs elsewhere","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Query set, capture, crawl, metrics, hash chain and sealing (no model; runs on CPU)","name":"decosa-api geo-audit (decosa_api/verticals/geo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge (mention, same-name confusion, position, sentiment, recommendation, brands) and the open-model answer panel","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Query set, capture, crawl, metrics, hash chain and sealing (no model; runs on CPU)","name":"decosa-api geo-audit (decosa_api/verticals/geo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge (mention, same-name confusion, position, sentiment, recommendation, brands) and the open-model answer panel","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."}]},{"id":"best","label":"Best · two 96 GB cards, a second open panel","gpu_gb":249.6,"basis":"stack","unknown":[],"components":[{"id":"runner","role":"Query set, capture, crawl, metrics, hash chain and sealing (no model; runs on CPU)","name":"decosa-api geo-audit (decosa_api/verticals/geo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge (mention, same-name confusion, position, sentiment, recommendation, brands) and the open-model answer panel","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"panel2","role":"Second open-model answer panel (planned)","name":"DeepSeek-V4-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · a four-family answer panel","gpu_gb":812.2,"basis":"estimate","unknown":[],"components":[{"id":"runner","role":"Query set, capture, crawl, metrics, hash chain and sealing (no model; runs on CPU)","name":"decosa-api geo-audit (decosa_api/verticals/geo)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"judge","role":"Judge (mention, same-name confusion, position, sentiment, recommendation, brands) and the open-model answer panel","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'."},{"id":"panel2","role":"Second open-model answer panel (planned)","name":"DeepSeek-V4-Flash (NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Third answer panel","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."},{"id":"nemotron-ultra-wanted","role":"Fourth answer panel","name":"Nemotron-3-Ultra-550B-A55B","where":"gpu","load":"resident","gb":370.6,"min_gb":370.6,"weights_gb":308,"basis":"estimate","precision":"fp4","gpus":1,"source":"Estimate: 550B parameters at about 4.5 bits per weight is about 308 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":{"fit":"full","memory_gb":32,"tier":"standard"},"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 192 GB)."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 96 GB)."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (32 of 64 GB)."}}},{"id":"review-reply","name":"Review reply with patient privacy","page":"https://decosa.ai/tools/operations/review-reply#self-host","tiers":[{"id":"lite","label":"Lite · one 48 GB card","gpu_gb":40,"basis":"estimate","unknown":[],"components":[{"id":"llm-lite","role":"Would draft the reply and run the privacy check; not measured on this task.","name":"Gemma 4 26B A4B (instruction-tuned)","where":"gpu","load":"resident","gb":40,"min_gb":32,"weights_gb":26,"basis":"estimate","precision":null,"gpus":1,"source":"Gemma 4 26B A4B, FP8 at load: BF16 weights are 49 GB (stack.json); quantised to FP8 at load time they are about half. The stacks name a 48 GB card for this, not measured."}]},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"llm","role":"Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).","name":"Qwen3.8-27B (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)"}]},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","gpu_gb":192,"basis":"stack","unknown":[],"components":[{"id":"llm-best","role":"Would draft the reply and run the privacy check; not measured on this task.","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."}]},{"id":"wanted","label":"Wanted · two large judges from different families","gpu_gb":384,"basis":"estimate","unknown":[],"components":[{"id":"llm-best","role":"Would draft the reply and run the privacy check; not measured on this task.","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","where":"gpu","load":"resident","gb":192,"min_gb":180,"weights_gb":176,"basis":"stack","precision":"fp4","gpus":2,"source":"DeepSeek-V4-Flash NVFP4: About 159-176 GB of weights (stack.json), run with tensor parallel 2 across two 96 GB cards on our server with little room left. The 180 GB minimum is an estimate."},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","where":"gpu","load":"resident","gb":192,"min_gb":185,"weights_gb":180,"basis":"estimate","precision":"fp4","gpus":2,"source":"GLM-5.3-Flash NVFP4: The stacks say it fits two 96 GB cards (321B parameters at about 4.5 bits is about 180 GB of weights: an estimate). Its sm_120 runtime does not work on our server today, so it has not been run here."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVIDIA NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)"},"rtx-5090x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size)."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (57.6 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"best","reason":"The standard tier fits (57.6 of 192 GB). The best tier fits too."},"m3-ultra-96x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."},"m5-max-64x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon."}}},{"id":"what-studies-found","name":"What studies found","page":"https://decosa.ai/tools/life-sciences/what-studies-found#self-host","tiers":[{"id":"lite","label":"Lite · one 32 GB card, no reranker","gpu_gb":57.6,"basis":"stack","unknown":[],"components":[{"id":"engine","role":"Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)","name":"decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstract","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"}]},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","gpu_gb":68.2,"basis":"estimate","unknown":[],"components":[{"id":"engine","role":"Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)","name":"decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)","where":"cpu","load":"resident","gb":null,"min_gb":null,"weights_gb":null,"basis":null,"precision":null,"gpus":1,"source":"Runs on CPU (vram_gb 0 in stack.json)."},{"id":"llm","role":"Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstract","name":"Qwen3.8-27B (NVFP4)","where":"gpu","load":"resident","gb":57.6,"min_gb":28,"weights_gb":21.4,"basis":"stack","precision":"fp4","gpus":1,"source":"Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)"},{"id":"reranker","role":"Ranks the PubMed results by relevance to the supplement and outcome before any abstract is read","name":"Qwen3-Reranker-4B (evidence retrieval block)","where":"gpu","load":"resident","gb":10.6,"min_gb":10.6,"weights_gb":8,"basis":"estimate","precision":null,"gpus":1,"source":"Estimate: 4B parameters at 2 bytes (BF16) per weight is about 8 GB, plus 20% working memory and 1 GB of runtime. Not measured."}]}],"mac":null,"undetermined":[],"presets":{"cpu-64x1":{"verdict":"no","recommended":null,"reason":"Qwen3.8-27B (NVFP4) needs a GPU."},"rtx-4090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 30.6 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes."},"rtx-5090x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier does not fit: Needs about 38.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes."},"rtx-5090x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions."},"l40sx1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"h100x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU."},"rtx-pro-6000x1":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (68.2 of 96 GB)."},"rtx-pro-6000x2":{"verdict":"runs","recommended":"standard","reason":"The standard tier fits (68.2 of 192 GB)."},"m3-ultra-96x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Qwen3-Reranker-4B (evidence retrieval block) has no mapped Apple Silicon build The lite tier fits with changes."},"m5-max-64x1":{"verdict":"smaller","recommended":"lite","reason":"The standard tier can't be checked: Qwen3-Reranker-4B (evidence retrieval block) has no mapped Apple Silicon build The lite tier fits with changes."}}}]}