The best Mac for local LLMs is the one with the most unified memory you can afford, on the fastest memory bus for that money. In 2026 that means a Mac mini M6 to start, a Mac mini M5 Pro for value, a Mac Studio M5 Max for comfortable 70B-class models, and an M5 Ultra for frontier-scale open weights.
Jump to: The two numbers · How much memory · Bandwidth ladder · What to buy · 128–512 GB · Traps · Used · Mac vs PC · FAQ
The two numbers that decide everything
On a PC your model must fit in GPU VRAM, and consumer cards top out at 24–32 GB unless you spend thousands. Apple Silicon uses unified memory: the GPU shares the whole memory pool with the CPU, so a Mac with 64 GB can load models that would need a multi-GPU PC rig. Two numbers decide what you get:
- How much unified memory — it determines which models you can load at all.
- Memory bandwidth (GB/s) — it determines how fast they generate. Decoding reads the active weights from memory for every token, so tokens per second scales almost linearly with bandwidth.
CPU core count and Neural Engine specs barely matter for decoding speed, so don't pay for them. The M5 and M6 GPUs add a Neural Accelerator to every core, which mainly speeds up prompt processing; see the FAQ below.
You only get part of your memory for models
macOS lets the GPU use about 75% of unified memory for a model and its KV cache; the rest stays with the system. That is the same figure the hardware hub and the analyzer use, so a "16 GB Mac" is a 12 GB model budget and a "64 GB Mac" is 48 GB. That limit is also why a Mac never refuses an oversized model outright: it swaps to disk instead, which looks like a freeze. Buying Intel second-hand is the other trap: read what an Intel Mac can and cannot run before you save money on one.
How much unified memory do you need?
The table below is computed by the same engine as the hardware hub and the VRAM calculator, so it cannot drift from them. The "entry Mac" column links to the machine's own page.
| Unified memory | Usable for models | What runs well (Q4_K_M) | Entry Mac that reaches it | Verdict |
|---|---|---|---|---|
| 16 GB | 12.0 GB | Cosmos 3 Nano (11.9 GB), Gemma 3 12B Instruct (11.1 GB) | Mac mini M6 16 GB | Entry level: chat and email |
| 24 GB | 18.0 GB | Mistral Small 3 (24B) (17.2 GB), Magistral Small 24B (17.0 GB) | Mac mini M6 24 GB | Best budget sweet spot |
| 32 GB | 24.0 GB | Qwen 3.5 35B-A3B (23.8 GB), Nex-N2.5 mini (23.8 GB) | Mac mini M6 32 GB | Comfortable daily driver |
| 36 GB | 27.0 GB | Gemma 3 27B Instruct (25.4 GB), K2 Horizon MoVA 36B-A4B (25.0 GB) | Mac Studio M5 Max 36 GB | Solid dev machine |
| 48 GB | 36.0 GB | Command R (35B) (32.7 GB), Gemma 3 27B Instruct (25.4 GB) | Mac mini M5 Pro 48 GB | The enthusiast sweet spot |
| 64 GB | 48.0 GB | Qwen 2.5 VL 72B Instruct (47.8 GB), Qwen 2.5 72B Instruct (47.0 GB) | Mac mini M5 Pro 64 GB | Serious local AI |
| 96 GB | 72.0 GB | GPT-OSS 120B (71.9 GB), Llama 4.5 Scout (69.4 GB) | Mac Studio M5 Ultra 96 GB | 70B-class with real context |
| 128 GB | 96.0 GB | Devstral-2 123B (77.9 GB), Qwen 3.5 122B-A10B (77.3 GB) | Mac Studio M5 Max 128 GB | Workstation class |
| 192 GB | 144 GB | MiniMax M2.7 230B-A10B (143 GB), MiniMax M2.5 230B (143 GB) | Mac Studio M5 Ultra 256 GB | Large MoE territory |
| 256 GB | 192 GB | MiMo-V2.5 310B (192 GB), DeepSeek V4-Flash (173 GB) | Mac Studio M5 Ultra 256 GB | Frontier-class open weights |
| 512 GB | 384 GB | Qwen3-Coder 480B-A35B (MoE) (292 GB), MiniMax M3 428B-A23B (264 GB) | Mac Studio M5 Ultra 512 GB (announced) | The largest Mac memory configuration |
Sized by the shared VRAM engine: total parameters at Q4_K_M, 8K context, no offload. Two practical notes: quantization is what makes these fit (see Quantization Explained), and long context eats gigabytes on top of the weights, so if you do RAG or long documents, buy one tier above your model size. You can check any specific model on any Mac in the hardware analyzer.
M4 vs M5 vs M6: the bandwidth ladder
Within a chip generation, the tier (base, Pro, Max, Ultra) matters far more than the generation itself, because the tier sets the memory bus width. The speeds below are the same modelled ranges the hardware hub shows for each chip.
| Chip | Bandwidth | Llama 3.1 8B Q4_K_M (estimated) |
|---|---|---|
| Apple M2 Pro | 200 GB/s | 13–26 tok/s |
| Apple M3 Pro | 153 GB/s | 9.7–20 tok/s |
| Apple M2 Max | 400 GB/s | 24–49 tok/s |
| Apple M4 | 120 GB/s | 7.7–16 tok/s |
| Apple M4 Pro | 273 GB/s | 17–35 tok/s |
| Apple M4 Max | 546 GB/s | 31–64 tok/s |
| Apple M5 | 153 GB/s | 9.7–20 tok/s |
| Apple M5 Pro | 307 GB/s | 19–39 tok/s |
| Apple M5 Max | 614 GB/s | 34–71 tok/s |
| Apple M6 | 170 GB/s | 11–22 tok/s |
| Apple M5 Ultra | 1229 GB/s | 59–123 tok/s |
Speeds are the same modelled ranges the hardware hub shows for each chip — a bandwidth-roofline estimate, not a benchmark. Read the bandwidth column first. The M6 is Apple's first 2 nm chip and a modest step up from the M5 at the base tier, but it still sits far below any Pro or Max: a base chip with a lot of memory can hold a big model and then generate slowly. The Pro, Max and Ultra tiers are where decode speed actually changes.
What to buy, by budget
The table lists every pick with its unified memory, bandwidth and Apple's launch price where we have a corroborated one. Each card below it is bound to that exact machine, so the memory in the heading and the memory on the card cannot disagree. Browse everything Apple sells in the hardware hub, the laptops view and the AI stations view.
| Mac | Unified memory | Bandwidth | Launch price | Status |
|---|---|---|---|---|
| Mac mini M6 16 GB | 16 GB unified | 170 GB/s | $899 | Available |
| Mac mini M6 24 GB | 24 GB unified | 170 GB/s | — | Available |
| Mac mini M6 32 GB | 32 GB unified | 170 GB/s | — | Available |
| MacBook Air 13 M5 24 GB | 24 GB unified | 153 GB/s | — | Available |
| Mac mini M5 Pro 24 GB | 24 GB unified | 307 GB/s | $1,699 | Available |
| Mac mini M5 Pro 48 GB | 48 GB unified | 307 GB/s | — | Available |
| Mac mini M5 Pro 64 GB | 64 GB unified | 307 GB/s | — | Available |
| Mac Studio M5 Max 36 GB | 36 GB unified | 460 GB/s | $2,499 | Available |
| Mac Studio M5 Max 128 GB | 128 GB unified | 614 GB/s | — | Available |
| Mac Studio M5 Ultra 96 GB | 96 GB unified | 1229 GB/s | — | Available |
| Mac Studio M5 Ultra 256 GB | 256 GB unified | 1229 GB/s | — | Available |
| Mac Studio M5 Ultra 512 GB | 512 GB unified | 1229 GB/s | — | Announced |
"—" means no corroborated launch price is recorded for that configuration yet; check Apple's configurator. Prices are Apple's launch price for the base configuration, not a current retail price.
Entry: Mac mini M6, 16 GB
The cheapest current entry: 2 nm, 170 GB/s, silent and always on. Be plain about the limit: 16 GB caps you at 8B-class models with a modest context. Fine for chat, summaries and a coding assistant on a small model; not for anything bigger.
The value step: Mac mini M6, 24 GB
The 24 GB configuration is the value inflection for the base chip: it opens up the 14–27B class where models stop feeling like toys. The 32 GB option adds headroom for context but not speed, since the bus stays the same.
Laptop: MacBook Air M5
Fanless and light, with the same bandwidth as any base M5. Honest framing: it holds mid-size models in 24 or 32 GB, but sustained generation heats a fanless chassis, and you pay laptop prices for base-chip bandwidth. Buy it because you need a laptop, not because it is a good LLM value.
Best value: Mac mini M5 Pro, 24–64 GB
The Pro chip roughly doubles the base bandwidth and reaches 64 GB in a small always-on box. For most people running local models every day, this is the best value Apple sells: pair it with our Personal AI Home Server guide and it serves your phone and laptop too. Go to 48 GB or more if you want 32B-class models with full context.
70B comfort: Mac Studio M5 Max
The Max tier is where 70B-class models become comfortable, and the 128 GB configuration hosts several models at once. For a desk-bound machine the Studio gives you the same silicon as the MacBook Pro without the battery and screen.
Frontier: Mac Studio M5 Ultra
The Ultra tier has the highest bandwidth Apple sells and the largest memory pools. The next section covers what that unlocks.
The frontier tier: 128–512 GB
At 128 GB and above, capacity stops being the thing you count and starts being the thing you spend. The models that live here are large open-weight Mixture-of-Experts models: only a few experts are read per token, so bandwidth is less punishing than the parameter count suggests, but every expert must stay resident, so total size decides whether a model loads at all. That is the VRAM-versus-bandwidth split the whole site is built on.
The table names models from our own model library and lets the fit engine decide each verdict. A "No" is a real "No": for example, the largest trillion-parameter models do not fit even in 512 GB at Q4_K_M.
| Model | Needs (Q4_K_M, 8K) | 128 GB | 192 GB | 256 GB | 512 GB |
|---|---|---|---|---|---|
| GPT-OSS 120B | 71.9 GB | Fits | Fits | Fits | Fits |
| Qwen 3.5 122B-A10B | 77.3 GB | Fits | Fits | Fits | Fits |
| Qwen 3 235B-A22B (MoE) | 144 GB | No | No | Fits | Fits |
| DeepSeek V4-Flash | 173 GB | No | No | Tight | Fits |
| Qwen 3.5 397B-A17B | 245 GB | No | No | No | Fits |
| Qwen3-Coder 480B-A35B (MoE) | 292 GB | No | No | No | Fits |
| DeepSeek V3.2 671B | 415 GB | No | No | No | No |
| Kimi K2.6 | 605 GB | No | No | No | No |
Each tier is judged against 75% of its unified memory — the same rule as everywhere else on the site. MoE models need their TOTAL parameters resident even though only the active experts are read per token. Three things to know about this tier:
- The 512 GB configuration is announced, not shipping. Apple dates the Mac Studio M5 Ultra with 512 GB to late October 2026 and press reports say orders may slip. Treat it as pre-release until it is on sale.
- Clustering is Apple's claim, not our measurement. Apple says Thunderbolt 5 and RDMA let you cluster Mac Studios into a shared memory pool, and that four clustered machines deliver up to 3x the AI inference of one (Apple newsroom). We have not benchmarked it.
- Compare against renting first. A 512 GB Mac Studio only pays off if you run it a lot. Our local vs cloud cost calculator and the cost tool put your usage against rented GPU hours.
The bandwidth traps
The M3 Pro is slower for LLMs than the older M2 Pro. Apple cut the memory bus that generation; the ladder above shows both figures. And a base chip with a large memory option is never a good LLM buy: for local AI, a used M2 Max beats a new base chip every time. If you are choosing between the memory upgrade and the next chip tier, choose the chip tier.
Buying used: the value picks
A used M1 Max or M2 Max Mac Studio with 64 GB (400 GB/s) regularly sells for $1,100–1,400 and runs 70B at Q4, which makes it the cheapest legitimate 70B machine you can buy. A used M4 Pro Mac mini is still a very good local-AI buy now that the M6 and M5 Pro have replaced it. Apple Silicon has no mining-abuse history, and unified memory cannot degrade like a GPU's fans and thermal pads, so used Macs are far lower-risk than used GPUs (compare our used GPU guide).
Mac vs a PC build
A PC with an RTX 5090 (32 GB, 1,792 GB/s) generates several times faster than any Mac for models that fit in 32 GB. The Mac's advantage starts exactly where VRAM ends: at 70B and beyond, the PC needs more GPUs (see Multi-GPU Inference) while the Mac just needs the memory it already has, silently and at a fraction of the power. Speed per dollar: PC. Model size per dollar and per watt: Mac.
Frequently asked questions
How much RAM do I need for local AI on a Mac?
16 GB is the working minimum for 8B models, 24–32 GB is the value sweet spot for 14–27B, 48 GB covers the 32B class, and 64 GB or more is where 70B models become practical. Buy one tier above the biggest model you plan to run, because context and multitasking need room.
Is a base chip like the M6 good enough for local LLMs?
For small models, yes. The M6 runs 8B-class models acceptably and, at 24 GB or more, opens the 14–27B class. But its bus is far narrower than any Pro or Max, so larger models generate slowly whatever the memory. Step up to a Pro chip before adding RAM.
Is the M6 Mac mini good enough for local AI?
Yes, as an entry machine. It is the cheapest current Mac, has 2 nm efficiency and up to 32 GB of unified memory. Its limit is bandwidth: it suits chat, summaries and small coding models, not 70B work. Choose 24 GB or 32 GB over 16 GB.
M6 or M5 Pro: which Mac mini for LLMs?
The M5 Pro, if you can afford it. It has roughly double the M6's memory bandwidth and reaches 64 GB, so it runs larger models faster. The M6 is the better buy only if your models fit in 24–32 GB and your budget is tight.
Do I need a Mac Studio, or is a Mac mini enough?
A Mac mini M5 Pro handles models up to about the 32B–70B range if you buy 48–64 GB. Choose a Mac Studio when you need the Max or Ultra bandwidth, more than 64 GB of memory, or several large models running at once.
Can a Mac really run a 70B model?
Yes. A Mac with 64 GB of unified memory or more runs Llama 3.3 70B at Q4 comfortably; the tier table above shows what each memory size holds. On 48 GB a 70B model technically fits only at a lower quantization with little room for context.
How much does a Mac that runs a 70B model cost in 2026?
The cheapest new route is a 64 GB Mac mini M5 Pro; see its hardware page for Apple's price. The cheapest route overall is a used M1 or M2 Max Mac Studio with 64 GB, which sells for roughly $1,100–1,400 in mid-2026 listings.
Can a Mac run gpt-oss-120b or a frontier open-weight model?
gpt-oss-120b fits in 96 GB of usable memory, so a 128 GB Mac runs it. Larger MoE models need 256 GB or 512 GB; the frontier table above gives each verdict, and the biggest trillion-parameter models do not fit even in 512 GB at Q4_K_M.
Is the MacBook Air M5 usable for local LLMs?
Yes for small and mid-size models, with caveats. It has base-chip bandwidth and no fan, so sustained generation throttles. Buy it if you need a light laptop that can run 8B–14B models; buy a MacBook Pro or a desktop for heavier work.
Do the Neural Accelerators in M5 and M6 GPU cores speed up LLM inference?
Mostly prompt processing, not generation. Apple's per-core Neural Accelerators speed up the compute-bound prefill stage, so long prompts start answering sooner. Token-by-token decoding is limited by memory bandwidth, which is why the bandwidth ladder still predicts generation speed.
Should I wait for the next chip or buy now?
Bandwidth, not generation, drives LLM speed, and used or refurbished high-bandwidth Macs are already excellent value. If a current configuration fits your model tier and budget, waiting rarely pays. The exception is the 512 GB Mac Studio, which is announced but not yet shipping.