Best Local LLMs for the Apple M2 Ultra
Written by Jakub Rusinowski · Last updated July 12, 2026
Best all-round pick: Qwen 3 235B-A22B (MoE)
The Apple M2 Ultra has 192 GB of unified memory, of which about 144 GB is available to a model. The largest model it can hold is Qwen 3 235B-A22B (MoE) (235B, 142.7 GB at Q4_K_M).
Best models overall on the Apple M2 Ultra
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| Qwen 3 235B-A22B (MoE) | 97 | 142.7 GB | 125K | Apache 2.0 |
| GPT-OSS 120B | 94.9 | 71.3 GB | 128K | Apache-2.0 |
| MiniMax M2.5 230B | 94.2 | 139.6 GB | 977K | Modified MIT (attribution required) |
| MiniMax M3-VL | 93.8 | 139.7 GB | 977K | Modified MIT (attribution required) |
| Qwen 3.5 122B-A10B (MoE) | 93.3 | 74.5 GB | 125K | Apache 2.0 |
| Qwen 3.5 122B-A10B | 91.7 | 74.5 GB | 128K | Apache-2.0 |
Best model by what you are doing
| Workload | Recommended on this GPU |
|---|---|
| Coding | Devstral-2 123B (98.2) MiniMax M2.5 230B (97.4) Qwen 3 235B-A22B (MoE) (96.8) |
| General assistant | Qwen 3 235B-A22B (MoE) (97) GPT-OSS 120B (94.9) MiniMax M2.5 230B (94.2) |
| Reasoning | Qwen 3 235B-A22B (MoE) (100) GLM-5.1 72B (98) Qwen 3.5 122B-A10B (MoE) (96.4) |
| RAG | MiniMax M2.5 230B (96.2) MiniMax M3-VL (95.7) Llama 4 Scout 17B (93.1) |
| Agents | Devstral-2 123B (94.5) Qwen3.8-Flash-Next (93.9) Qwen3-Coder-Next (80B-A3B MoE) (92.7) |
| Vision | MiniMax M3-VL (99.4) Qwen 3.5 72B (96.4) Llama 3.2 90B Vision Instruct (94.7) |
How these numbers are calculated
- Usable memory is 144 GB of the card's 192 GB, after the share macOS reserves for the system.
- Fit is measured at Q4_K_M — quantized weights plus KV cache plus 0.8 GB runtime overhead.
- Within this card's budget, models are ranked on how well they use the memory available, not on being small — so the recommendation scales with the hardware.
FAQ
What is the best LLM for the Apple M2 Ultra?
Qwen 3 235B-A22B (MoE) is the strongest all-round pick that fits its 144 GB.
What is the largest model the Apple M2 Ultra can run?
Qwen 3 235B-A22B (MoE) — 235B parameters, needing 142.7 GB at Q4_K_M.
How much of the Apple M2 Ultra's memory can a model actually use?
About 144 GB of its 192 GB, because macOS reserves a share of unified memory for the system.
Compatibility Checks for the Apple M2 Ultra
- Can I run DeepSeek V3 on the Apple M2 Ultra?
- Can I run Qwen 3 on the Apple M2 Ultra?
- Can I run Qwen 3.5 on the Apple M2 Ultra?
- Can I run MiniMax M2.5 on the Apple M2 Ultra?
- Can I run DeepSeek V3.2 on the Apple M2 Ultra?
By Workload
- Best local LLMs for coding
- Best local LLMs for general assistant
- Best local LLMs for reasoning
- Best local LLMs for rag