AppleMac / unified memoryCurrent

Apple M5 Max (32-core GPU) for local LLMs

Written by Jakub Rusinowski · Last updated

With 36 GB at 460 GB/s, the M5 Max (32-core GPU) runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 27B Instruct (~25.4 GB), and the top pick is Qwen 3.6 35B-A3B at 37–77 tok/s.

The 18-core CPU / 32-core GPU M5 Max: 460 GB/s against the 40-core part's 614 GB/s, and a 36 GB configuration. Modelling a 36 GB machine on the 40-core bandwidth overstates every speed by a third. The 36 GB ceiling for this bin is not yet independently confirmed.

Models that run on the M5 Max (32-core GPU)

Q4_K_M, 8K context, 27 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Qwen 3.6 35B-A3B
Qwen 3.6
21.9 GB37–77 tok/s
Nex-N2.5 mini
Nex-N2.5
21.9 GB37–77 tok/s
Nex-N2 mini
Nex-N2
21.9 GB37–77 tok/s
Qwen 3.5 35B-A3B
Qwen 3.5
21.9 GB37–77 tok/s
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB37–77 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB20–42 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB35–73 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB46–96 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB46–96 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB38–78 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB38–78 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB38–78 tok/s
Showing 12 of 109

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA DGX Spark (128GB) — 128 GB VRAM · 150 W board powerDeploy in the cloud now NVIDIA A40 on RunPod — from $0.44/hr · rate checked 2026-08

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-09-30.

Memory
36 GB
Memory bandwidth
460 GB/s
Architecture
ARM, 3nm TSMC
Series
Apple Silicon
Board power
35 W
Release year
2026
Compute backends
METAL
Usable for models
27 GB (~75% of unified)
Best forMac Studio base32B modelsMacBook Pro

Similar GPUs

Frequently asked questions

Can the Apple M5 Max (32-core GPU) run local LLMs?

Yes. With 36 GB (27 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 27B Instruct, needing about 25.4 GB.

How fast is the Apple M5 Max (32-core GPU) for AI inference?

It is estimated to run Llama 3.1 8B at 27–56 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 27 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 36 GB?

Among the best that fit: Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B, Laguna XS 2.1 33B-A3B.