Local LLM Hardware Heatmap

Green = best  |  Red = worst  |  Per-column OKLCH palette gradient  |  Click any header to sort  |  Budget cap: $20,000

SystemSystem Cost (k$)Memory (GB)Bandwidth (GB/s)Power (W)t/s (32B)t/k$ (32B)t/kW (32B)t/s (70B)t/k$ (70B)t/kW (70B)t/s (104B)t/k$ (104B)t/kW (104B)
Mac Studio M5 Ultra 96GB5.5961,200300448.00146.7254.5583.3162.9153.3
Mac Studio M5 Ultra 256GB9.52561,200300444.63146.7252.6383.3161.6853.3
Mac Studio M5 Max 128GB4.0128614160358.75218.8205.00125.071.7543.8
Mac Studio M3 Ultra 96GB5.396819280305.66107.1173.2160.7112.0839.3
Mac Studio M4 Max 128GB3.5128546180205.72111.192.5750.061.7133.3
Mac Mini Pro2.06420085126.00141.242.0047.121.0023.5
Dual RTX 40906.4481,0081,2006510.1654.2264.0621.71.50.231.2
Triple RTX 30905.5729361,400407.2728.6203.6414.3122.188.6
Quad 4060 Ti 16GB3.464288800185.3022.582.3610.041.185.0
Dual RTX 30904.2489369004510.7150.0184.2920.01.50.361.7
Single RTX 50906.6321,792750558.3373.3304.5540.01.50.232.0
Dual RTX 509011.6643,5841,350907.7666.7332.8424.480.695.9
Ryzen AI Max+ 3953.6128270120154.11125.051.3741.73.50.9629.2
Single RTX 40904.0241,0087505012.5066.741.005.310.251.3
Tiiny AI Pocket Lab2.08020565147.00215.463.0092.34.52.2569.2
NVIDIA DGX Spark4.7128273170183.83105.951.0629.43.50.7420.6
Quadro RTX 8000 48GB11.648672260282.41107.7110.9542.31.80.166.9
RTX PRO 6000 Blackwell 96GB14.1961,790800553.9068.8322.2740.0100.7112.5
AMD Radeon Pro W79005.148864295316.08105.112.52.4542.42.00.396.8
AMD Instinct MI210 64GB7.3641,600300446.03146.7212.8870.05.00.6916.7
AMD Radeon Pro W78004.832576260224.5684.640.8315.40.80.173.1
AMD Radeon Pro Duo 32GB2.63244825062.3124.010.384.00.20.080.8

Methodology

Budget cap: All systems are normalized to complete acquisition cost under a $20,000 ceiling (September 2026 pricing), shown in k$ (thousands USD). The system cost (k$) column is normalized to full local-build/system cost, not raw add-in card MSRP. Discrete workstation GPUs were converted to system pricing with a fixed $1,600 host allowance, while passive datacenter accelerators use a $3,000 server-platform allowance.

September 2026 market refresh (verified street/used pricing): RTX 5090 32GB new street ~$5,000 (Sep 1, 2026: cheapest US listings $4,930–$5,170, 2.5× the $1,999 MSRP amid the GDDR7 shortage; Tom's Hardware tracker low $4,799), so Single = $5,000 + $1,600 host = $6,600 and Dual = $10,000 + $1,600 = $11,600. Used RTX 4090 ~$2,400 (eBay sold-listing average $2,362 on Sep 5, 2026, range $2,000–$2,541). Used RTX 3090 ~$1,300 (Aug 2026 eBay median $1,275, up ~50% since January on local-AI demand). RTX 4060 Ti 16GB new ~$449. AMD Radeon PRO W7900 ~$3,500 (Sep 2026 lowest-average $3,157); W7800 ~$3,229; used MI210 ~$4,299 + $3,000 server platform. Strix Halo 128GB boxes (Ryzen AI Max+ 395, GMKtec EVO-X2) ~$3,649 — DRAM shortage roughly doubled 128GB mini-PC prices since launch. Apple announced the Mac Studio M5 Max ($2,499 base; 128GB config ~$3,999) and M5 Ultra ($5,499 base with 96GB, +$4,000 for 256GB; 1.2 TB/s bandwidth, shipping Sept 22) on Aug 25, 2026, discontinuing the M4 Max and M3 Ultra; the M5 Ultra 512GB (late October, expected well above $10k) is not yet priced. M5 Max/M5 Ultra throughput is estimated from the 614 GB/s / 1.2 TB/s bandwidth (M5 Max: ~110 t/s 8B, ~20 t/s 70B; M5 Ultra: ~90 t/s 8B at Q4). RTX PRO 6000 Blackwell Workstation 96GB card ~$12,500 (NVIDIA marketplace $13,250 after the Aug 13 raise toward $16,000; street spans $10,499 B&H to $16,999 Newegg, used $9,500–$11,000), so system = $12,500 + $1,600 = $14,100. NVIDIA DGX Spark $4,699. Quadro RTX 8000 48GB reflects new-old-stock pricing.

Vendor-sourced fields: VRAM, memory bandwidth, and board power use vendor-published specs where available: RTX 5090 (32GB, 1,792 GB/s, 575W), RTX PRO 6000 Blackwell (96GB ECC GDDR7, 1,790 GB/s, 600W), Mac Studio M3 Ultra (96GB unified, 819 GB/s), and AMD product pages for W7800, W7900, and MI210.

Estimated fields: LLM throughput and ratio values are model-side inference estimates (Q4_K_M llama.cpp class) based on memory capacity, memory bandwidth, architecture class, and software maturity — anchored to community 2026 benchmarks (RTX 5090: ~145 t/s 8B, ~55 t/s 32B, ~38 t/s 70B tight-fit; Dual RTX 5090: ~33 t/s 70B Q4 tensor-parallel; RTX PRO 6000: ~32 t/s 70B, 96GB holds 120B-class) rather than direct apples-to-apples runs on identical hosts. The 70B figure on the Single RTX 5090 reflects a near-full 32GB fit with minimal context headroom.

Ratio columns (grouped at the end): The nine t/* columns are the cross product of the three reference models {32B, 70B, 104B} and the three units {t/s, t/k$, t/kW}. t/s = raw throughput; t/k$ = t/s per $1,000 of system cost; t/kW = t/s per kilowatt of total system power.