Apple M4 Max (32-core GPU) for local LLMs
Written by Jakub Rusinowski · Last updated
With 36 GB at 410 GB/s, the M4 Max (32-core GPU) runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 27B Instruct (~25.4 GB), and the top pick is Qwen 3.6 35B-A3B at 33–70 tok/s.
The 14-core CPU / 32-core GPU M4 Max: 410 GB/s rather than the 40-core part's 546 GB/s, and a 36 GB configuration. Quoting the faster bin at a 36 GB machine's owner overstates every speed on the page.
Models that run on the M4 Max (32-core GPU)
Q4_K_M, 8K context, 27 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 27 GB | VRAM | Speed |
|---|---|---|---|
| Qwen 3.6 35B-A3B Qwen 3.6 | 21.9 GB | 33–70 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 33–70 tok/s | |
| Nex-N2 mini Nex-N2 | 21.9 GB | 33–70 tok/s | |
| Qwen 3.5 35B-A3B Qwen 3.5 | 21.9 GB | 33–70 tok/s | |
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | 34–70 tok/s | |
| Granite 4.0 Small-H 32B-A9B IBM Granite 4.0 | 20.1 GB | 18–38 tok/s | |
| Nemotron-Cascade 2 30B-A3B Nemotron Cascade 2 | 19.9 GB | 32–66 tok/s | |
| Qwen 3 30B-A3B (MoE) Qwen 3 | 20 GB | 42–87 tok/s | |
| Qwen3-Coder 30B-A3B (MoE) Qwen3-Coder | 20 GB | 42–87 tok/s | |
| GLM-4.7-Flash 30B-A3B GLM-4.7 / GLM-Z1 | 18.9 GB | 34–71 tok/s | |
| Nemotron 3 Nano Omni 30B-A3B Nemotron 3 Nano Omni | 18.9 GB | 34–71 tok/s | |
| North Mini Code 1.0 30B-A3B North Mini Code | 18.9 GB | 34–71 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-09-30.
- Memory
- 36 GB
- Memory bandwidth
- 410 GB/s
- Architecture
- ARM, 3nm TSMC
- Series
- Apple Silicon
- Board power
- 35 W
- Release year
- 2024
- Compute backends
- METAL
- Usable for models
- 27 GB (~75% of unified)
Similar GPUs
Frequently asked questions
Can the Apple M4 Max (32-core GPU) run local LLMs?
Yes. With 36 GB (27 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 27B Instruct, needing about 25.4 GB.
How fast is the Apple M4 Max (32-core GPU) for AI inference?
It is estimated to run Llama 3.1 8B at 24–50 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 27 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 36 GB?
Among the best that fit: Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B, Laguna XS 2.1 33B-A3B.