Local AI on macOS
Written by Jakub Rusinowski · Last updated September 30, 2026
Unified memory lets a Mac load models far larger than any consumer GPU can hold, at a fraction of the memory bandwidth — capacity you can buy, speed you cannot.
Runtimes that work on macOS
| Runtime | How to install | Acceleration | Watch out for |
|---|---|---|---|
| Ollama | Official macOS app, or Homebrew.brew install ollama | Metal (Apple Silicon) | — |
| LM Studio | Official macOS app; ships both the Metal and MLX runtimes. | Metal, MLX | — |
| MLX | Apple's own array framework, installed via pip.pip install mlx-lm | Metal (Apple Silicon) | Usually the fastest option on Apple Silicon, but needs MLX-format weights — not every model has an MLX conversion. |
| llama.cpp | Homebrew, or build from source (Metal is enabled by default on Apple Silicon).brew install llama.cpp | Metal, CPU | — |
What macOS is good at
- The cheapest way to get 64–256 GB of accelerator-addressable memory in one box.
- Very low idle power and near-silent operation, which makes an always-on local assistant practical.
- MLX gives Apple Silicon a first-party, actively developed inference path.
macOS limitations
- macOS caps how much unified memory the GPU may claim; the default ceiling is roughly 75% of total memory, which is why usable figures on this site are below the sticker capacity.
- Memory is soldered on every Apple Silicon Mac — the configuration you buy is permanent, so buy the memory you will need later.
- Memory bandwidth is well below a discrete GPU at the same price: a Mac that loads a 70B model still generates far slower than a GPU that can hold it.
- No CUDA, so tooling that assumes CUDA — including most fine-tuning stacks — will not run.
When it breaks on macOS
- The Mac swaps to disk instead of refusing an oversized model
- Raising the Metal working-set ceiling
- Gatekeeper says the app is damaged
- What actually runs on an Intel Mac
- MLX will not install, or runs on the CPU
- Fast for two minutes, then half as fast — thermal throttling
Hardware and models on macOS
macOS supports Apple accelerators — 23 of the 109 in our hardware database, up to 512 GB.
| Memory | Hardware on this platform | Models to run |
|---|---|---|
| 8 GB | — | Qwen 3 8B GLM-6 9B GLM-4 9B |
| 16 GB | Apple M1 | Mistral Small 3.1 24B Qwen 3.5 14B Qwen 3.5 14B |
| 24 GB | Apple M4 Pro Apple M2 Apple M3 | Gemma 4 27B ⭐ Mistral Small 3.1 24B Qwen3.8 27B |
| 48 GB | — | GLM-5.1 72B Qwen 3.5 72B Gemma 4 27B ⭐ |
| 96 GB | Apple M3 Max (30-core GPU) Apple M2 Max | GPT-OSS 120B Qwen 3.5 122B-A10B (MoE) GLM-5.1 72B |
| 192 GB | Apple M2 Ultra | Qwen 3 235B-A22B (MoE) DeepSeek V4.1 Flash DeepSeek V4-Flash |
Machines that run macOS
| Machine | Memory for models | Form factor | Upgradeable |
|---|---|---|---|
| Mac Studio (M5 Ultra, 512 GB) | 512 GB unified | Desktop | No |
| Mac Studio (M5 Ultra, 256 GB) | 256 GB unified | Desktop | No |
| MacBook Pro 16" (M4 Max, 128 GB) | 128 GB unified | Laptop | No |
| Mac Studio (M5 Ultra, 96 GB) | 96 GB unified | Desktop | No |
| MacBook Pro 16" (M4 Max, 48 GB) | 48 GB unified | Laptop | No |
| Mac Studio (M3 Ultra, 256 GB) | 256 GB unified | Desktop | No |
| Mac Studio (M5 Max 40-core GPU, 128 GB) | 128 GB unified | Desktop | No |
| MacBook Pro 14" (M4 Pro, 24 GB) | 24 GB unified | Laptop | No |
FAQ
What is the best way to run an LLM locally on macOS?
Ollama — Official macOS app, or Homebrew. It reaches Metal (Apple Silicon).
Which GPUs work for local AI on macOS?
Apple hardware, up to 512 GB in our database.
What are the downsides of running local AI on macOS?
macOS caps how much unified memory the GPU may claim; the default ceiling is roughly 75% of total memory, which is why usable figures on this site are below the sticker capacity.
Machines Running This Platform
Hardware for This Platform
- Best models for the Apple M5 Ultra
- Best models for the Apple M3 Ultra
- Best models for the Apple M2 Ultra
- Best models for the Apple M4 Max