Apple M2 Max for local LLMs
Written by Jakub Rusinowski · Last updated
With 96 GB at 400 GB/s, the M2 Max runs 126 catalogued models at Q4_K_M with 8K context. The largest that fits is GPT-OSS 120B (~71.9 GB), and the top pick is Qwen3-Coder-Next (80B-A3B MoE) at 29–60 tok/s.
Up to 96 GB unified memory, enough for 70B-class models at Q4_K_M. Popular in Mac Studio configurations. Silent, 40W system power. A strong used value at $1,500–2,000.
Models that run on the M2 Max
Q4_K_M, 8K context, 72 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 72 GB | VRAM | Speed |
|---|---|---|---|
| Qwen3-Coder-Next (80B-A3B MoE) Qwen3-Coder | 49.1 GB | 29–60 tok/s | |
| Kolibri 1 Kolibri | 48.4 GB | 49–102 tok/s | |
| K2 Horizon MoVA 36B-A4B K2 Horizon | 25 GB | 30–63 tok/s | |
| Qwen 3.6 35B-A3B Qwen 3.6 | 22.1 GB | 54–112 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 33–68 tok/s | |
| Nex-N2 mini Nex-N2 | 21.9 GB | 33–68 tok/s | |
| Qwen 3.5 35B-A3B Qwen 3.5 | 21.9 GB | 33–68 tok/s | |
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | 33–69 tok/s | |
| Granite 4.0 Small-H 32B-A9B IBM Granite 4.0 | 20.1 GB | 18–37 tok/s | |
| Nemotron-Cascade 2 30B-A3B Nemotron Cascade 2 | 19.9 GB | 31–65 tok/s | |
| Qwen 3 30B-A3B (MoE) Qwen 3 | 20 GB | 41–86 tok/s | |
| Qwen3-Coder 30B-A3B (MoE) Qwen3-Coder | 20 GB | 41–86 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-09-29.
- Memory
- 96 GB
- Memory bandwidth
- 400 GB/s
- Architecture
- ARM, 5nm TSMC
- Series
- Apple Silicon
- Board power
- 40 W
- Release year
- 2023
- Usable for models
- 72 GB (~75% of unified)
Similar GPUs
Frequently asked questions
Can the Apple M2 Max run local LLMs?
Yes. With 96 GB (72 GB usable by a model) it runs 126 of the catalogued models at Q4_K_M with 8K context; the largest is GPT-OSS 120B, needing about 71.9 GB.
How fast is the Apple M2 Max for AI inference?
It is estimated to run Llama 3.1 8B at 24–49 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 3.2–6.7 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 96 GB?
Among the best that fit: Qwen3-Coder-Next (80B-A3B MoE), Kolibri 1, K2 Horizon MoVA 36B-A4B, Qwen 3.6 35B-A3B, Nex-N2.5 mini. The quickest start is Ollama: ollama run llama3.3:70b.
Keep going
Can I run it on the M2 Max?
More for this card
Tools