AppleMac / unified memoryCurrent

Apple M1 for local LLMs

Written by Jakub Rusinowski · Last updated

With 16 GB at 68 GB/s, the M1 runs 69 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 12B Instruct (~11.1 GB), and the top pick is LFM2.5-8B-A1B at 11–23 tok/s.

The original Apple Silicon chip. Up to 16 GB unified memory at 68 GB/s limits it to smaller 7–8B models in Q4. The fanless MacBook Air still works well for lightweight local AI.

Models that run on the M1

Q4_K_M, 8K context, 12 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
LFM2.5-8B-A1B
LFM2.5
5.8 GB11–23 tok/s
Gemma 3n E2B
Gemma 3n
4.1 GB11–22 tok/s
Cosmos 3 Edge
Cosmos 3
3.2 GB11–23 tok/s
DeepSeek-OCR 3B
DeepSeek-OCR
2.6 GB19–39 tok/s
EXAONE 3.5 2.4B
EXAONE 3.5
2.9 GB11–23 tok/s
Qwen 3.5 2B
Qwen 3.5
2 GB12–24 tok/s
VibeThinker 1.5B
VibeThinker
1.7 GB14–29 tok/s
MiniCPM-V 4.6
MiniCPM-V
1.6 GB16–33 tok/s
Llama 3.2 1B Instruct
Llama 3.2 Family
1.8 GB21–45 tok/s
Qwen 3.5 0.8B
Qwen 3.5
1.3 GB21–44 tok/s
SmolLM2 360M Instruct
SmolLM2
1.4 GB36–76 tok/s
SmolLM2 1.7B Instruct
SmolLM2
3.4 GB8.8–18 tok/s
Showing 12 of 69

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Apple Mac mini M1 (16GB)
16 GB VRAM · 20 W board power
2026 prices are volatile — check the current listing.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-07-12.

Memory
16 GB
Memory bandwidth
68 GB/s
Architecture
ARM, 5nm TSMC
Series
Apple Silicon
Board power
20 W
Release year
2020
Usable for models
12 GB (~75% of unified)
Best for8B models (Q4)MacBook Airentry-level Apple Silicon
Recommended first model (Ollama)

Similar GPUs

Frequently asked questions

Can the Apple M1 run local LLMs?

Yes. With 16 GB (12 GB usable by a model) it runs 69 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 12B Instruct, needing about 11.1 GB.

How fast is the Apple M1 for AI inference?

It is estimated to run Llama 3.1 8B at 4.4–9.2 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 12 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 16 GB?

Among the best that fit: LFM2.5-8B-A1B, Gemma 3n E2B, Cosmos 3 Edge, DeepSeek-OCR 3B, EXAONE 3.5 2.4B. The quickest start is Ollama: ollama run llama3.2:3b.