Apple M1 Max for local LLMs
Written by Jakub Rusinowski · Last updated
With 64 GB at 400 GB/s, the M1 Max runs 118 catalogued models at Q4_K_M with 8K context. The largest that fits is Llama 3.3 70B Instruct (~45.7 GB), and the top pick is K2 Horizon MoVA 36B-A4B at 30–63 tok/s.
The original Apple Silicon powerhouse — up to 64 GB unified memory, about 48 GB usable by a model: enough for 70B-class models at Q4_K_M with little headroom. Still excellent in 2025. Available as used MacBook Pro M1 Max for $800–1,200.
Models that run on the M1 Max
Q4_K_M, 8K context, 48 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 48 GB | VRAM | Speed |
|---|---|---|---|
| K2 Horizon MoVA 36B-A4B K2 Horizon | 25 GB | 30–63 tok/s | |
| Qwen 3.6 35B-A3B Qwen 3.6 | 22.1 GB | 54–112 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 33–68 tok/s | |
| Nex-N2 mini Nex-N2 | 21.9 GB | 33–68 tok/s | |
| Qwen 3.5 35B-A3B Qwen 3.5 | 21.9 GB | 33–68 tok/s | |
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | 33–69 tok/s | |
| Granite 4.0 Small-H 32B-A9B IBM Granite 4.0 | 20.1 GB | 18–37 tok/s | |
| Nemotron-Cascade 2 30B-A3B Nemotron Cascade 2 | 19.9 GB | 31–65 tok/s | |
| Qwen 3 30B-A3B (MoE) Qwen 3 | 20 GB | 41–86 tok/s | |
| Qwen3-Coder 30B-A3B (MoE) Qwen3-Coder | 20 GB | 41–86 tok/s | |
| GLM-4.7-Flash 30B-A3B GLM-4.7 / GLM-Z1 | 18.9 GB | 34–70 tok/s | |
| Nemotron 3 Nano Omni 30B-A3B Nemotron 3 Nano Omni | 18.9 GB | 34–70 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.77/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-09-29.
- Memory
- 64 GB
- Memory bandwidth
- 400 GB/s
- Architecture
- ARM, 5nm TSMC
- Series
- Apple Silicon
- Board power
- 30 W
- Release year
- 2021
- Usable for models
- 48 GB (~75% of unified)
Similar GPUs
Frequently asked questions
Can the Apple M1 Max run local LLMs?
Yes. With 64 GB (48 GB usable by a model) it runs 118 of the catalogued models at Q4_K_M with 8K context; the largest is Llama 3.3 70B Instruct, needing about 45.7 GB.
How fast is the Apple M1 Max for AI inference?
It is estimated to run Llama 3.1 8B at 24–49 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 3.2–6.7 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 64 GB?
Among the best that fit: K2 Horizon MoVA 36B-A4B, Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B. The quickest start is Ollama: ollama run llama3.3:70b.
Keep going
Can I run it on the M1 Max?
More for this card
Tools