AppleMac / unified memoryCurrent

Apple M1 Max for local LLMs

Written by Jakub Rusinowski · Last updated

With 64 GB at 400 GB/s, the M1 Max runs 118 catalogued models at Q4_K_M with 8K context. The largest that fits is Llama 3.3 70B Instruct (~45.7 GB), and the top pick is K2 Horizon MoVA 36B-A4B at 30–63 tok/s.

The original Apple Silicon powerhouse — up to 64 GB unified memory, about 48 GB usable by a model: enough for 70B-class models at Q4_K_M with little headroom. Still excellent in 2025. Available as used MacBook Pro M1 Max for $800–1,200.

Models that run on the M1 Max

Q4_K_M, 8K context, 48 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
K2 Horizon MoVA 36B-A4B
K2 Horizon
25 GB30–63 tok/s
Qwen 3.6 35B-A3B
Qwen 3.6
22.1 GB54–112 tok/s
Nex-N2.5 mini
Nex-N2.5
21.9 GB33–68 tok/s
Nex-N2 mini
Nex-N2
21.9 GB33–68 tok/s
Qwen 3.5 35B-A3B
Qwen 3.5
21.9 GB33–68 tok/s
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB33–69 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB18–37 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB31–65 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB41–86 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB41–86 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB34–70 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB34–70 tok/s
Showing 12 of 118

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Apple MacBook Pro M1 Max
64 GB VRAM · 30 W board power
2026 prices are volatile — check the current listing.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-09-29.

Memory
64 GB
Memory bandwidth
400 GB/s
Architecture
ARM, 5nm TSMC
Series
Apple Silicon
Board power
30 W
Release year
2021
Usable for models
48 GB (~75% of unified)
Best for70B modelsMacBook Pro originalpower efficiency
Recommended first model (Ollama)

Similar GPUs

Frequently asked questions

Can the Apple M1 Max run local LLMs?

Yes. With 64 GB (48 GB usable by a model) it runs 118 of the catalogued models at Q4_K_M with 8K context; the largest is Llama 3.3 70B Instruct, needing about 45.7 GB.

How fast is the Apple M1 Max for AI inference?

It is estimated to run Llama 3.1 8B at 24–49 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 3.2–6.7 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 64 GB?

Among the best that fit: K2 Horizon MoVA 36B-A4B, Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B. The quickest start is Ollama: ollama run llama3.3:70b.