AMD Ryzen AI Max+ 395 — Local LLM Performance & Compatibility
作者: Jakub Rusinowski · 最后更新: 2026年7月12日
"Strix Halo" SoC pairing 16 Zen 5 CPU cores with a 40-CU RDNA 3.5 iGPU (Radeon 8060S), sharing up to 128 GB of LPDDR5X-8000 memory at 256 GB/s. Up to 96 GB can be allocated to the GPU as VRAM via AMD Variable Graphics Memory (Windows) or GTT (Linux) — the figure used here. Powers AI mini-PCs like the Framework Desktop, GMKtec EVO-X2, and HP Z2 Mini G1a. ROCm support is improving but still behind CUDA.
Technical Specifications
| VRAM | 128 GB |
| Memory Bandwidth | 256 GB/s |
| TDP | 120 W |
| Architecture | Zen 5 + RDNA 3.5 "Strix Halo" |
| Release Year | 2025 |
| MSRP at Launch | $1,999 |
| Inference Speed (Llama 3.1 8B Q4_K_M) | 18–38 tok/s (estimated) |
| Inference Speed (Llama 3.3 70B Q4_K_M) | 2.4–5.1 tok/s (estimated) |
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
LLMs Compatible with 128 GB VRAM
All models below run comfortably in 128 GB VRAM with Q4_K_M quantization.
| Qwen3.8 | Qwen3.8-Flash-Next · 109 GB VRAM · Q4_K_M · qwen3-8 |
| Qwen 3.5 | Qwen 3.5 122B-A10B · 74 GB VRAM · Q4_K_M · ollama run qwen3.5:122b |
| Mistral Small 4 | Mistral Small 4 119B-A6.5B · 73 GB VRAM · Q4_K_M · ollama run mistral-small |
| Llama 4 | Llama 4 Scout 17B · 67 GB VRAM · Q4_K_M · ollama run llama4:scout |
| Command R Family | Command R+ (104B) · 64 GB VRAM · Q4_K_M · ollama run command-r-plus |
| Llama 3.2 Family | Llama 3.2 90B Vision Instruct · 54 GB VRAM · Q4_K_M · llama-3-2 |
| Llama 3.2 Vision | Llama 3.2 Vision 90B · 54 GB VRAM · Q4_K_M · ollama run llama3.2-vision:90b |
| Qwen3-Coder | Qwen3-Coder 80B-A3B (MoE) · 49 GB VRAM · Q4_K_M · ollama run qwen3-coder:80b-a3b-q4 |
62 more families also fit 128 GB — browse the full model library.
Best Use Cases
- 70B models
- AMD AI mini-PC
- unified memory
- ROCm
Quick Start with Ollama
Install Ollama then run the recommended model for this GPU:
ollama run llama3.3:70b
FAQ
Can the AMD Ryzen AI Max+ 395 run local LLMs?
Yes — the AMD Ryzen AI Max+ 395 has 128 GB VRAM and runs "Strix Halo" SoC pairing 16 Zen 5 CPU cores with a 40-CU RDNA 3.5 iGPU (Radeon 8060S), sharing up to 128 GB of LPDDR5X-8
How fast is the AMD Ryzen AI Max+ 395 for AI inference?
The AMD Ryzen AI Max+ 395 is estimated to run Llama 3.1 8B at 18–38 tok/s with Q4_K_M quantization. For Llama 3.3 70B the estimate is 2.4–5.1 tok/s. These are modelled estimates, not measurements — see /en/methodology.
What LLMs can I run on 128 GB VRAM?
With 128 GB you can run: Qwen3.8, Qwen 3.5, Mistral Small 4, Llama 4, Command R Family. Use Ollama for the easiest setup: ollama run llama3.3:70b.
Can I Run It? — AMD Ryzen AI Max+ 395
- Nemotron 70B on AMD Ryzen AI Max+ 395
- Command R Family on AMD Ryzen AI Max+ 395
- Qwen 2.5 Family on AMD Ryzen AI Max+ 395
- Llama 4 on AMD Ryzen AI Max+ 395
- Llama 3.2 Family on AMD Ryzen AI Max+ 395
- Qwen 2.5 VL on AMD Ryzen AI Max+ 395
- Qwen3-Coder on AMD Ryzen AI Max+ 395
- Llama 3.2 Vision on AMD Ryzen AI Max+ 395
Compare Similar GPUs
- NVIDIA GB10 Grace Blackwell (128 GB, 0 t/s)
- Apple M4 Max (128 GB, 110 t/s)
- Apple M3 Max (128 GB, 95 t/s)
- Apple M1 Ultra (128 GB, 155 t/s)
VRAM Tier
Buying a laptop?
Buying Guide
← All GPU Reviews | All Hardware | Check Your Hardware | Full Benchmarks | Can I Run It?