Best Local LLMs for the Apple M2 Ultra

Written by Jakub Rusinowski · Last updated July 12, 2026

Best all-round pick: Qwen 3 235B-A22B (MoE)

The Apple M2 Ultra has 192 GB of unified memory, of which about 144 GB is available to a model. The largest model it can hold is Qwen 3 235B-A22B (MoE) (235B, 142.7 GB at Q4_K_M).

Best models overall on the Apple M2 Ultra

ModelScoreMemory at Q4_K_MContextLicence
Qwen 3 235B-A22B (MoE)97142.7 GB125KApache 2.0
GPT-oss 120B95.173.3 GB125KApache-2.0
MiniMax M3 230B-A10B94.4139.7 GB977KModified MIT (attribution required)
MiniMax M2.5 230B94.2139.7 GB977KModified MIT (attribution required)
Nemotron 70B Instruct94.143.4 GB125KLlama Community
Qwen 3.5 122B-A10B (MoE)93.374.5 GB125KApache 2.0

Best model by what you are doing

WorkloadRecommended on this GPU
CodingDevstral-2 123B (97.8)
MiniMax M3 230B-A10B (97.7)
MiniMax M2.5 230B (97.4)
General assistantQwen 3 235B-A22B (MoE) (97)
GPT-oss 120B (95.1)
MiniMax M3 230B-A10B (94.4)
ReasoningQwen 3 235B-A22B (MoE) (100)
GLM-5.1 72B (98)
Qwen 3.5 122B-A10B (MoE) (96.4)
RAGMiniMax M3 230B-A10B (96.3)
MiniMax M2.5 230B (96.2)
Llama 4 Scout 17B (93.1)
AgentsMiniMax M2.7 230B-A10B (89.1)
Ternary Bonsai 27B (88.8)
Qwen 3.5 122B-A10B (MoE) (86.9)
VisionMiniMax M3-VL (99.4)
Qwen 3.5 72B (96.4)
Llama 3.2 90B Vision Instruct (94.7)

How these numbers are calculated

FAQ

What is the best LLM for the Apple M2 Ultra?

Qwen 3 235B-A22B (MoE) is the strongest all-round pick that fits its 144 GB.

What is the largest model the Apple M2 Ultra can run?

Qwen 3 235B-A22B (MoE) — 235B parameters, needing 142.7 GB at Q4_K_M.

How much of the Apple M2 Ultra's memory can a model actually use?

About 144 GB of its 192 GB, because macOS reserves a share of unified memory for the system.

Compatibility Checks for the Apple M2 Ultra

By Workload

More

← All GPUs | Apple M2 Ultra specs