Apple M1 for local LLMs
Written by Jakub Rusinowski · Last updated
With 16 GB at 68 GB/s, the M1 runs 69 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 12B Instruct (~11.1 GB), and the top pick is LFM2.5-8B-A1B at 11–23 tok/s.
The original Apple Silicon chip. Up to 16 GB unified memory at 68 GB/s limits it to smaller 7–8B models in Q4. The fanless MacBook Air still works well for lightweight local AI.
Models that run on the M1
Q4_K_M, 8K context, 12 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 12 GB | VRAM | Speed |
|---|---|---|---|
| LFM2.5-8B-A1B LFM2.5 | 5.8 GB | 11–23 tok/s | |
| Gemma 3n E2B Gemma 3n | 4.1 GB | 11–22 tok/s | |
| Cosmos 3 Edge Cosmos 3 | 3.2 GB | 11–23 tok/s | |
| DeepSeek-OCR 3B DeepSeek-OCR | 2.6 GB | 19–39 tok/s | |
| EXAONE 3.5 2.4B EXAONE 3.5 | 2.9 GB | 11–23 tok/s | |
| Qwen 3.5 2B Qwen 3.5 | 2 GB | 12–24 tok/s | |
| VibeThinker 1.5B VibeThinker | 1.7 GB | 14–29 tok/s | |
| MiniCPM-V 4.6 MiniCPM-V | 1.6 GB | 16–33 tok/s | |
| Llama 3.2 1B Instruct Llama 3.2 Family | 1.8 GB | 21–45 tok/s | |
| Qwen 3.5 0.8B Qwen 3.5 | 1.3 GB | 21–44 tok/s | |
| SmolLM2 360M Instruct SmolLM2 | 1.4 GB | 36–76 tok/s | |
| SmolLM2 1.7B Instruct SmolLM2 | 3.4 GB | 8.8–18 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-07-12.
- Memory
- 16 GB
- Memory bandwidth
- 68 GB/s
- Architecture
- ARM, 5nm TSMC
- Series
- Apple Silicon
- Board power
- 20 W
- Release year
- 2020
- Usable for models
- 12 GB (~75% of unified)
Similar GPUs
Frequently asked questions
Can the Apple M1 run local LLMs?
Yes. With 16 GB (12 GB usable by a model) it runs 69 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 12B Instruct, needing about 11.1 GB.
How fast is the Apple M1 for AI inference?
It is estimated to run Llama 3.1 8B at 4.4–9.2 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 12 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 16 GB?
Among the best that fit: LFM2.5-8B-A1B, Gemma 3n E2B, Cosmos 3 Edge, DeepSeek-OCR 3B, EXAONE 3.5 2.4B. The quickest start is Ollama: ollama run llama3.2:3b.
Keep going
Can I run it on the M1?
More for this card
Tools