Green = best | Red = worst | Per-column OKLCH palette gradient | Click any header to sort | Budget cap: $20,000
| System | System Cost (k$) | Memory (GB) | Bandwidth (GB/s) | Power (W) | t/s (32B) | t/k$ (32B) | t/kW (32B) | t/s (70B) | t/k$ (70B) | t/kW (70B) | t/s (104B) | t/k$ (104B) | t/kW (104B) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mac Studio M5 Ultra 96GB | 5.5 | 96 | 1,200 | 300 | 44 | 8.00 | 146.7 | 25 | 4.55 | 83.3 | 16 | 2.91 | 53.3 |
| Mac Studio M5 Ultra 256GB | 9.5 | 256 | 1,200 | 300 | 44 | 4.63 | 146.7 | 25 | 2.63 | 83.3 | 16 | 1.68 | 53.3 |
| Mac Studio M5 Max 128GB | 4.0 | 128 | 614 | 160 | 35 | 8.75 | 218.8 | 20 | 5.00 | 125.0 | 7 | 1.75 | 43.8 |
| Mac Studio M3 Ultra 96GB | 5.3 | 96 | 819 | 280 | 30 | 5.66 | 107.1 | 17 | 3.21 | 60.7 | 11 | 2.08 | 39.3 |
| Mac Studio M4 Max 128GB | 3.5 | 128 | 546 | 180 | 20 | 5.72 | 111.1 | 9 | 2.57 | 50.0 | 6 | 1.71 | 33.3 |
| Mac Mini Pro | 2.0 | 64 | 200 | 85 | 12 | 6.00 | 141.2 | 4 | 2.00 | 47.1 | 2 | 1.00 | 23.5 |
| Dual RTX 4090 | 6.4 | 48 | 1,008 | 1,200 | 65 | 10.16 | 54.2 | 26 | 4.06 | 21.7 | 1.5 | 0.23 | 1.2 |
| Triple RTX 3090 | 5.5 | 72 | 936 | 1,400 | 40 | 7.27 | 28.6 | 20 | 3.64 | 14.3 | 12 | 2.18 | 8.6 |
| Quad 4060 Ti 16GB | 3.4 | 64 | 288 | 800 | 18 | 5.30 | 22.5 | 8 | 2.36 | 10.0 | 4 | 1.18 | 5.0 |
| Dual RTX 3090 | 4.2 | 48 | 936 | 900 | 45 | 10.71 | 50.0 | 18 | 4.29 | 20.0 | 1.5 | 0.36 | 1.7 |
| Single RTX 5090 | 6.6 | 32 | 1,792 | 750 | 55 | 8.33 | 73.3 | 30 | 4.55 | 40.0 | 1.5 | 0.23 | 2.0 |
| Dual RTX 5090 | 11.6 | 64 | 3,584 | 1,350 | 90 | 7.76 | 66.7 | 33 | 2.84 | 24.4 | 8 | 0.69 | 5.9 |
| Ryzen AI Max+ 395 | 3.6 | 128 | 270 | 120 | 15 | 4.11 | 125.0 | 5 | 1.37 | 41.7 | 3.5 | 0.96 | 29.2 |
| Single RTX 4090 | 4.0 | 24 | 1,008 | 750 | 50 | 12.50 | 66.7 | 4 | 1.00 | 5.3 | 1 | 0.25 | 1.3 |
| Tiiny AI Pocket Lab | 2.0 | 80 | 205 | 65 | 14 | 7.00 | 215.4 | 6 | 3.00 | 92.3 | 4.5 | 2.25 | 69.2 |
| NVIDIA DGX Spark | 4.7 | 128 | 273 | 170 | 18 | 3.83 | 105.9 | 5 | 1.06 | 29.4 | 3.5 | 0.74 | 20.6 |
| Quadro RTX 8000 48GB | 11.6 | 48 | 672 | 260 | 28 | 2.41 | 107.7 | 11 | 0.95 | 42.3 | 1.8 | 0.16 | 6.9 |
| RTX PRO 6000 Blackwell 96GB | 14.1 | 96 | 1,790 | 800 | 55 | 3.90 | 68.8 | 32 | 2.27 | 40.0 | 10 | 0.71 | 12.5 |
| AMD Radeon Pro W7900 | 5.1 | 48 | 864 | 295 | 31 | 6.08 | 105.1 | 12.5 | 2.45 | 42.4 | 2.0 | 0.39 | 6.8 |
| AMD Instinct MI210 64GB | 7.3 | 64 | 1,600 | 300 | 44 | 6.03 | 146.7 | 21 | 2.88 | 70.0 | 5.0 | 0.69 | 16.7 |
| AMD Radeon Pro W7800 | 4.8 | 32 | 576 | 260 | 22 | 4.56 | 84.6 | 4 | 0.83 | 15.4 | 0.8 | 0.17 | 3.1 |
| AMD Radeon Pro Duo 32GB | 2.6 | 32 | 448 | 250 | 6 | 2.31 | 24.0 | 1 | 0.38 | 4.0 | 0.2 | 0.08 | 0.8 |
Budget cap: All systems are normalized to complete acquisition cost under a $20,000 ceiling (September 2026 pricing), shown in k$ (thousands USD). The system cost (k$) column is normalized to full local-build/system cost, not raw add-in card MSRP. Discrete workstation GPUs were converted to system pricing with a fixed $1,600 host allowance, while passive datacenter accelerators use a $3,000 server-platform allowance.
September 2026 market refresh (verified street/used pricing): RTX 5090 32GB new street ~$5,000 (Sep 1, 2026: cheapest US listings $4,930–$5,170, 2.5× the $1,999 MSRP amid the GDDR7 shortage; Tom's Hardware tracker low $4,799), so Single = $5,000 + $1,600 host = $6,600 and Dual = $10,000 + $1,600 = $11,600. Used RTX 4090 ~$2,400 (eBay sold-listing average $2,362 on Sep 5, 2026, range $2,000–$2,541). Used RTX 3090 ~$1,300 (Aug 2026 eBay median $1,275, up ~50% since January on local-AI demand). RTX 4060 Ti 16GB new ~$449. AMD Radeon PRO W7900 ~$3,500 (Sep 2026 lowest-average $3,157); W7800 ~$3,229; used MI210 ~$4,299 + $3,000 server platform. Strix Halo 128GB boxes (Ryzen AI Max+ 395, GMKtec EVO-X2) ~$3,649 — DRAM shortage roughly doubled 128GB mini-PC prices since launch. Apple announced the Mac Studio M5 Max ($2,499 base; 128GB config ~$3,999) and M5 Ultra ($5,499 base with 96GB, +$4,000 for 256GB; 1.2 TB/s bandwidth, shipping Sept 22) on Aug 25, 2026, discontinuing the M4 Max and M3 Ultra; the M5 Ultra 512GB (late October, expected well above $10k) is not yet priced. M5 Max/M5 Ultra throughput is estimated from the 614 GB/s / 1.2 TB/s bandwidth (M5 Max: ~110 t/s 8B, ~20 t/s 70B; M5 Ultra: ~90 t/s 8B at Q4). RTX PRO 6000 Blackwell Workstation 96GB card ~$12,500 (NVIDIA marketplace $13,250 after the Aug 13 raise toward $16,000; street spans $10,499 B&H to $16,999 Newegg, used $9,500–$11,000), so system = $12,500 + $1,600 = $14,100. NVIDIA DGX Spark $4,699. Quadro RTX 8000 48GB reflects new-old-stock pricing.
Vendor-sourced fields: VRAM, memory bandwidth, and board power use vendor-published specs where available: RTX 5090 (32GB, 1,792 GB/s, 575W), RTX PRO 6000 Blackwell (96GB ECC GDDR7, 1,790 GB/s, 600W), Mac Studio M3 Ultra (96GB unified, 819 GB/s), and AMD product pages for W7800, W7900, and MI210.
Estimated fields: LLM throughput and ratio values are model-side inference estimates (Q4_K_M llama.cpp class) based on memory capacity, memory bandwidth, architecture class, and software maturity — anchored to community 2026 benchmarks (RTX 5090: ~145 t/s 8B, ~55 t/s 32B, ~38 t/s 70B tight-fit; Dual RTX 5090: ~33 t/s 70B Q4 tensor-parallel; RTX PRO 6000: ~32 t/s 70B, 96GB holds 120B-class) rather than direct apples-to-apples runs on identical hosts. The 70B figure on the Single RTX 5090 reflects a near-full 32GB fit with minimal context headroom.
Ratio columns (grouped at the end): The nine t/* columns are the cross product of the three reference models {32B, 70B, 104B} and the three units {t/s, t/k$, t/kW}. t/s = raw throughput; t/k$ = t/s per $1,000 of system cost; t/kW = t/s per kilowatt of total system power.