Written by Jakub Rusinowski · Last updated July 12, 2026
Best all-round pick: Qwen 3.5 397B
The Apple M3 Ultra has 512 GB of unified memory, of which about 384 GB is available to a model. The largest model it can hold is Qwen3-Coder 480B-A35B (MoE) (480B, 290.6 GB at Q4_K_M).
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| Qwen 3.5 397B | 96.3 | 240.5 GB | 125K | Apache 2.0 |
| Qwen 3.5 397B-A17B | 95.9 | 240.5 GB | 256K | Apache 2.0 |
| GLM-6 355B-A32B | 95.4 | 215.1 GB | 195K | MIT |
| Llama 4 Maverick 17B | 95.3 | 242.3 GB | 1M | Llama Community |
| Qwen 3 235B-A22B (MoE) | 93.8 | 142.7 GB | 125K | Apache 2.0 |
| DeepSeek V4.1 Flash | 93.8 | 172.3 GB | 977K | MIT |
| Workload | Recommended on this GPU |
|---|---|
| Coding | Qwen3-Coder 480B-A35B (MoE) (100) GLM-6 355B-A32B (99.1) DeepSeek V4.1 Flash (97.8) |
| General assistant | Qwen 3.5 397B (96.3) Qwen 3.5 397B-A17B (95.9) GLM-6 355B-A32B (95.4) |
| Reasoning | Qwen 3.5 397B (98.9) Llama 4 Maverick 17B (98.2) Qwen 3 235B-A22B (MoE) (98.1) |
| RAG | DeepSeek V4.1 Flash (97.6) DeepSeek V4-Flash (97.4) Qwen 3.5 397B-A17B (94.9) |
| Agents | Qwen3-Coder 480B-A35B (MoE) (97.2) GLM-6 355B-A32B (94.6) Ternary Bonsai 27B (87.8) |
| Vision | Qwen 3.5 397B (100) MiniMax M3-VL (96.8) Llama 3.2 90B Vision Instruct (91.5) |
Qwen 3.5 397B is the strongest all-round pick that fits its 384 GB.
Qwen3-Coder 480B-A35B (MoE) — 480B parameters, needing 290.6 GB at Q4_K_M.
About 384 GB of its 512 GB, because macOS reserves a share of unified memory for the system.