AppleMac / unified memoryCurrent

Apple M6 for local LLMs

Written by Jakub Rusinowski · Last updated

With 32 GB of LPDDR5X at 170 GB/s, the M6 runs 108 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B at 15–32 tok/s.

Apple's first 2 nm chip: 12-core CPU, 12-core GPU with a Neural Accelerator in every GPU core, up to 32 GB of unified memory at 170 GB/s (11% more than the M5). Ships in the Mac mini from 22 September 2026. The TDP figure is a placeholder pending a second source; it does not affect throughput or fit.

Models that run on the M6

Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB15–32 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB14–30 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB20–41 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB20–41 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB16–33 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB16–33 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB16–33 tok/s
Nemotron 3.5 Lightning 30B-A3B
Nemotron 3.5
18.9 GB16–33 tok/s
Gemma 4 26B-A4B
Gemma 4
16.5 GB14–29 tok/s
GPT-OSS 20B
GPT-OSS
13.8 GB21–44 tok/s
Gemma 4 E2B
Gemma 4
3.9 GB14–29 tok/s
Gemma 3n E2B
Gemma 3n
4.1 GB25–52 tok/s
Showing 12 of 108

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA GeForce RTX 5090 32GB — 32 GB VRAM · 575 W board powerDeploy in the cloud now NVIDIA A40 on RunPod — from $0.44/hr · rate checked 2026-08

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-09-30.

Memory
32 GB LPDDR5X
Memory bandwidth
170 GB/s
Architecture
ARM, 2nm TSMC
Series
Apple Silicon
Board power
25 W
Release year
2026
Compute backends
METAL
Usable for models
24 GB (~75% of unified)
Best for8B-27B modelsMac minialways-on home server

Similar GPUs

Frequently asked questions

Can the Apple M6 run local LLMs?

Yes. With 32 GB (24 GB usable by a model) it runs 108 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.

How fast is the Apple M6 for AI inference?

It is estimated to run Llama 3.1 8B at 11–22 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 32 GB?

Among the best that fit: Laguna XS 2.1 33B-A3B, Nemotron-Cascade 2 30B-A3B, Qwen 3 30B-A3B (MoE), Qwen3-Coder 30B-A3B (MoE), GLM-4.7-Flash 30B-A3B.