AI Stations and Mini-PCs for Local LLMs

Whole machines built around one large pool of memory. Capacity decides which models load at all; bandwidth decides whether they are usable — and on these boxes the two often point in opposite directions.

Specs verified 2026-07-21 · Prices verified 2026-07-06

DeviceMemoryBandwidthLlama 3.1 8B Q4_K_MPrice$/GBTDPReleasedRuns with
Mac Studio (M3 Ultra, 256 GB)256 GB unified819 GB/s43–90 tok/s60 W2025metal
Mac Studio (M2 Ultra, 192 GB)192 GB unified800 GB/s43–89 tok/s60 W2023metal
Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB unified256 GB/s18–38 tok/s120 W2025rocm, vulkan
NVIDIA DGX Spark128 GB unified273 GB/s26–53 tok/s$4,699$36.71150 W2025cuda, vulkan
AMD Ryzen AI Max+ 39596 GB unified256 GB/s18–38 tok/s$1,999$20.82120 W2025rocm, vulkan
Mac mini (M4 Pro, 64 GB)64 GB unified273 GB/s17–35 tok/s30 W2024metal
Mac Studio (M4 Max, 64 GB)64 GB unified546 GB/s31–64 tok/s35 W2025metal
Beelink SER9 (Ryzen AI 9, 32 GB)32 GB unified256 GB/s18–38 tok/s$859$26.84120 W2025rocm, vulkan
RTX 5090 Desktop (32 GB VRAM, 64 GB RAM)32 GB1,792 GB/s113–234 tok/s575 W2025cuda, vulkan
RTX 3090 Desktop (24 GB VRAM, 64 GB RAM)24 GB936 GB/s60–115 tok/s350 W2020cuda, vulkan
RTX 4090 Desktop (24 GB VRAM, 64 GB RAM)24 GB1,008 GB/s86–165 tok/s450 W2022cuda, vulkan
Mac mini (M4, 16 GB)16 GB unified120 GB/s7.7–16 tok/s$599$37.4422 W2024metal
RTX 5060 Ti 16 GB Desktop (16 GB VRAM, 32 GB RAM)16 GB448 GB/s40–83 tok/s180 W2025cuda, vulkan
RTX 5080 Desktop (16 GB VRAM, 32 GB RAM)16 GB960 GB/s74–153 tok/s360 W2025cuda, vulkan
RTX 3060 12 GB Desktop (12 GB VRAM, 32 GB RAM)12 GB360 GB/s26–50 tok/s170 W2021cuda, vulkan

Local AI Hardware — Every Device Compared · Best GPUs for Local AI — Full Comparison · Laptops for Local AI