Will it run on my hardware?
Every tool lists its models by quality tier. Pick your GPU, Mac or cloud box (or type in its memory) and the check packs those models onto it: which tier fits, what to swap (another precision, a 4-bit build, a smaller model) and a setup prompt with the plan written in.
Only two setups have been measured: an RTX PRO 6000 Blackwell 96 GB and a Mac Studio M3 Ultra. Everything else is a memory calculation, and every number says whether it is measured, stated in the tool's stack.json, or an estimate. Machine-readable: /api/hardware.json.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitVisit copilot on GeForce RTX 5090
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Lite · one 48 GB card, live pass only: what changesuses estimates
- Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
- Pass 1: Voxtral Mini 4B Realtime. ~24 GB (at least ~16 GB), weights 8.3 GB (from stack.json). Voxtral Mini 4B Realtime: Weights 8.3 GB in BF16; the compose file gives it 0.25 of a 96 GB card (24 GB) for streaming sessions (field stack.json). The field stack's lite tier puts it on a separate card of 16 GB or more.
- Language model for the lite tier: Qwen3.8-27B (official FP8). ~33.6 GB (at least ~32 GB), weights 29 GB (from stack.json). Qwen3.8-27B FP8: 33.6 GB is the sales lite tier's allotment (stack.json). Weights of about 29 GB are an estimate (27.8B parameters at one byte, plus higher-precision layers).
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Visit copilot, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Visit copilot on my hardware Fetch https://decosa.ai/prompts/clinical-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=clinical) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · one 48 GB card, live pass only (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Pass 1: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB - Language model for the lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB Warning: the fit check says this tier does not fit: Needs about 48 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/clinical-assemble.md
Hardware catalog
Memory, bandwidth and low-precision support decide what fits and how fast it decodes. NVFP4 builds need a Blackwell GPU; Hopper and Ada run FP8; Ampere runs FP8 checkpoints only weight-only; AMD runs FP8 or BF16 builds through ROCm (untested with these stacks); a Mac runs MLX builds in unified memory.
NVIDIA GeForce
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| GeForce RTX 3090 No FP8 tensor cores: FP8 checkpoints run weight-only (slower) in vLLM, NVFP4 ones do not run. | 24 GBGDDR6X | 936 GB/s | No | No | $1,499launch MSRP (2020). Sold used today; used prices not checked. | |
| GeForce RTX 4090 | 24 GBGDDR6X | 1,008 GB/s | Yes | No | $1,599launch MSRP (2022). Current street and used prices not checked. | |
| GeForce RTX 5090 | 32 GBGDDR7 | 1,792 GB/s | Yes | NVFP4 | $1,999launch MSRP (Jan 2025). Newegg listings on 2026-09-26 were $6,900-9,990, mostly marketplace sellers. |
NVIDIA workstation
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| RTX PRO 4000 Blackwell | 24 GBGDDR7 ECC | 672 GB/s | Yes | NVFP4 | not checked | |
| RTX PRO 4500 Blackwell | 32 GBGDDR7 ECC | 896 GB/s | Yes | NVFP4 | not checked | |
| RTX PRO 5000 Blackwell 48 GB | 48 GBGDDR7 ECC | 1,344 GB/s | Yes | NVFP4 | $8,599retail listing, 2026-09-26. Newegg: $8,599-10,450. | |
| RTX PRO 5000 Blackwell 72 GB | 72 GBGDDR7 ECC | 1,344 GB/s | Yes | NVFP4 | $9,999retail listing, 2026-09-26. Newegg: $9,999-17,468. | |
| RTX PRO 6000 Blackwell 96 GBmeasured here What Decosa runs: two of these in our server. Every measured GPU number on this site is from this card. | 96 GBGDDR7 ECC | 1,792 GB/s | Yes | NVFP4 | $17,999retail listing, 2026-09-26. Newegg: $17,999 (NVIDIA-branded), recertified from about $16,400. It launched at about $8,500 in 2025 (not confirmed from a fetched page). Max-Q (300 W) and Server Edition (1,597 GB/s) variants have the same 96 GB. | |
| RTX 6000 Ada Generation | 48 GBGDDR6 ECC | 960 GB/s* | Yes | No | $6,800*launch MSRP | |
| RTX A6000 Ampere: FP8 checkpoints run weight-only (slower) in vLLM; NVFP4 ones do not run. | 48 GBGDDR6 ECC | 768 GB/s* | No | No | $4,650*launch MSRP |
NVIDIA datacenter
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| L4 72 W card; low bandwidth, so large models decode slowly. | 24 GBGDDR6 | 300 GB/s | Yes | No | not checked | |
| L40S | 48 GBGDDR6 ECC | 864 GB/s | Yes | No | not checked | |
| A100 40 GB | 40 GBHBM2 | 1,555 GB/s* | No | No | not checked | |
| A100 80 GB 2,039 GB/s is the SXM figure; PCIe is 1,935 GB/s. | 80 GBHBM2e | 2,039 GB/s | No | No | not checked | |
| H100 80 GB (SXM) The H100 NVL has 94 GB at 3,900 GB/s; the PCIe card has 80 GB at about 2,000 GB/s. | 80 GBHBM3 | 3,350 GB/s | Yes | No | not checked | |
| H200 141 GB 4,800 GB/s is the SXM figure. | 141 GBHBM3e | 4,800 GB/s | Yes | No | not checked | |
| B200 180 GB Sold in 8-GPU HGX/DGX systems (1,440 GB in total). Datacentre Blackwell (sm_100): builds made for the RTX PRO 6000 (sm_120) may need a different engine image. | 180 GBHBM3e | 8,000 GB/s | Yes | NVFP4 | not checked | |
| B300 (Blackwell Ultra) NVIDIA's pages give about 262-278 GB per GPU (2.1 TB per 8 on HGX B300); 288 GB is often quoted. The fit check uses the lower figure. | 262 GBHBM3e* | 8,000 GB/s* | Yes | NVFP4 | not checked | |
| GB200 (per Blackwell GPU) Rack-scale: a GB200 superchip pairs one Grace CPU with two GPUs (372 GB); NVL72 is 72 GPUs, liquid-cooled, sold as a rack. Far beyond any single use case here. | 186 GBHBM3e | 8,000 GB/s | Yes | NVFP4 | not checked |
AMD
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| Instinct MI300X vLLM supports MI300 (gfx942) on ROCm. NVIDIA NVFP4 checkpoints do not run on AMD: use the FP8 or BF16 build. Decosa's containers are CUDA builds and have not been run on ROCm. | 192 GBHBM3 | 5,300 GB/s | Yes | No | not checked | |
| Instinct MI325X Same ROCm path as the MI300X; FP8 or BF16 checkpoints only. | 256 GBHBM3E* | 6,000 GB/s | Yes | No | not checked | |
| Instinct MI355X Native MXFP4 (AMD Quark checkpoints), not NVIDIA's NVFP4: the NVFP4 builds used here need an FP8 or MXFP4 replacement. vLLM needs ROCm 7.0 or newer for MI350-series cards. | 288 GBHBM3E | 8,000 GB/s | Yes | MXFP4 | not checked | |
| Radeon RX 7900 XTX vLLM lists RX 7900 (gfx1100) as supported on ROCm; quantisation support is narrower than on Instinct. Untested with Decosa's stacks. | 24 GBGDDR6 | 960 GB/s | No* | No | $999launch MSRP (2022) |
Apple Silicon (unified memory)
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| Apple M3 Max MacBook Pro (2023). 400 GB/s on the 16-core chip, 300 GB/s on the 14-core one. | 36 / 48 / 64 / 96 / 128 GBunified LPDDR5* | 400 GB/s | No | No | not checked | |
| Apple M4 Max MacBook Pro (2024) and Mac Studio (2025). 546 GB/s with the 40-core GPU, 410 GB/s with the 32-core one. Apple made no M4 Ultra. | 36 / 48 / 64 / 128 GBunified | 546 GB/s | No | No | not checked | |
| Apple M3 Ultra (Mac Studio)measured here The Mac every measured Mac number here comes from (512 GB). Replaced by the M5 Ultra Mac Studio in September 2026. | 96 / 256 / 512 GBunified | 819 GB/s | No | No | not checked | |
| Apple M5 MacBook Pro 14, MacBook Air, iPad Pro (from October 2025). | 16 / 24 / 32 GBunified* | 153 GB/s | No | No | not checked | |
| Apple M5 Pro MacBook Pro (March 2026) and Mac mini (September 2026). | 24 / 48 / 64 GBunified | 307 GB/s | No | No | not checked | |
| Apple M5 Max Mac Studio (September 2026) and MacBook Pro. 460 GB/s in the Mac Studio (32-core GPU); the 40-core GPU has 614 GB/s. | 36 / 48 / 64 / 128 GBunified | 460 GB/s | No | No | $2,499*Mac Studio starting price. 36 GB configuration; the 128 GB price was not checked. | |
| Apple M5 Ultra (Mac Studio) Mac Studio from September 2026. Not measured with Decosa's stacks yet. | 96 / 256 / 512 GBunified | 1,200 GB/s | No | No | $5,499*Mac Studio starting price. 96 GB configuration; the 512 GB price was not checked. |
CPU only
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| CPU only (no GPU) Rule engines, signing, OCR, embeddings and small voices run on CPU. Language models do not run at a usable speed here: point them at a remote endpoint instead. | 16 / 32 / 64 / 128 / 256 / 512 GBsystem RAM | n/a | No | No | not checked |
Cloud, rented by the hour
| Device | Memory | Bandwidth | FP8 | FP4 | Typical price | Sources |
|---|---|---|---|---|---|---|
| Cloud: H100 80 GB The H100 NVL has 94 GB at 3,900 GB/s; the PCIe card has 80 GB at about 2,000 GB/s. | 80 GBHBM3 | 3,350 GB/s | Yes | No | from $2.69/GPU-houron demand, per GPU-hour, 2026-09-26. RunPod $2.69-3.49 (SXM), Lambda $3.99-4.29. | |
| Cloud: H200 141 GB 4,800 GB/s is the SXM figure. | 141 GBHBM3e | 4,800 GB/s | Yes | No | from $3.59/GPU-houron demand, per GPU-hour, 2026-09-26. RunPod $3.59-4.59. | |
| Cloud: L40S 48 GB | 48 GBGDDR6 ECC | 864 GB/s | Yes | No | from $0.79/GPU-houron demand, per GPU-hour, 2026-09-26. RunPod $0.79-1.09. | |
| Cloud: RTX PRO 6000 96 GB What Decosa runs: two of these in our server. Every measured GPU number on this site is from this card. | 96 GBGDDR7 ECC | 1,792 GB/s | Yes | NVFP4 | from $1.69/GPU-houron demand, per GPU-hour, 2026-09-26. RunPod $1.69-2.09. The same card Decosa measures on. | |
| Cloud: B200 180 GB Sold in 8-GPU HGX/DGX systems (1,440 GB in total). Datacentre Blackwell (sm_100): builds made for the RTX PRO 6000 (sm_120) may need a different engine image. | 180 GBHBM3e | 8,000 GB/s | Yes | NVFP4 | from $5.98/GPU-houron demand, per GPU-hour, 2026-09-26. RunPod $5.98-6.79, Lambda $6.69-6.99. | |
| Cloud: MI300X 192 GB vLLM supports MI300 (gfx942) on ROCm. NVIDIA NVFP4 checkpoints do not run on AMD: use the FP8 or BF16 build. Decosa's containers are CUDA builds and have not been run on ROCm. | 192 GBHBM3 | 5,300 GB/s | Yes | No | from $2.99/GPU-houron demand, per GPU-hour, 2026-09-26. Hot Aisle $2.99 (VM), $3.39 (8x bare metal). |
* not confirmed from a fetched source on 2026-09-26. Prices are launch MSRPs or listings on that date; GPU street prices in 2026 are far above launch prices. Decosa's own box has 2x RTX PRO 6000 Blackwell 96 GB; the Mac numbers are from a Mac Studio M3 Ultra 512 GB (content/verticals/mac.json). The basis labels used on this page: measured, from stack.json, estimate.
How the check works
- Each model's memory is its weights plus KV cache plus runtime, as deployed. It comes from a footprint in
content/hardware.json(with its source), else the component'svram_gbin stack.json, else an estimate from parameters times bytes per weight plus 20%. - Servers that stay loaded (language and speech models) are packed onto the GPUs; renderers that load per job (image, video, music) count once, in what is left.
- If a model does not fit at its deployed size, the check tries its smallest setting (shorter context, fewer parallel sessions), then a listed replacement.
- A remote endpoint is offered as an option but never counted as fitting: the data would leave your machine.
- On a Mac, the standard tier uses the part-by-part mapping measured on a Mac Studio (/self-host/mac).
Components whose memory is not known yet (8)
- Check their brief: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (standard tier)
- CMMC / NIST 800-171 evidence map: A licence-clean document OCR and layout model (not chosen) (alternate-ocr-lane tier)
- Consented creator dubbing: InfiniteTalk (alternate-voice-lipsync tier)
- Decosa Studio: Z-Image-Turbo (lite tier)
- Family interview film: Decosa document reader (Docling layout + PaddleOCR-VL-1.6) (standard tier)
- Filing tie-out and MD&A grounding: A licence-clean table and document reader (not chosen) (alternate-docreader tier)
- Sanctions alert disposition record: A licence-clean multilingual transliteration or name-matching model (not chosen) (alternate-translit tier)
- SAR narrative desk: A licence-clean vision/OCR model (not chosen) (alternate-ocr tier)