Best Coding LLM for 8GB, 16GB & 24GB VRAM (2026)
On a 24 GB card (RTX 3090/4090), run Qwen 3.6 27B (~17.6 GB at Q4_K_M), or Devstral Small 2 24B (~15.3 GB) if you want an agentic specialist. With a 64 GB-class machine, Qwen3-Coder-Next (80B total, 3B active) is the most capable coder here that one machine can hold: all 80B parameters must fit, but only 3B are read per token, so it decodes fast. On 8 GB, use StarCoder 2 7B for autocomplete (~5.1 GB) — agentic work starts at 16 GB and is comfortable at 24 GB. Every VRAM figure below comes from our compute module, not marketing pages.
Written by Jakub Rusinowski · Last updated September 29, 2026 · Hardware figures computed by our VRAM engine
Most "best coding LLM" lists rank models nobody can run. This one is sorted the other way around: by the hardware tier you actually have. Each table shows the model's real Q4_K_M VRAM requirement — the same math behind our GPU compatibility checker — plus the GPU tier and the smallest Mac unified-memory config that fits it.
Two things to know before the tables. First, coding models split into agentic coders (built to drive an editing loop in Continue, Cline, or Aider) and autocomplete models (small, fill-in-the-middle-tuned, judged on latency). Serious local setups run one of each. Second, the frontier moves fast, and published coding benchmarks disagree between sources and scaffolds, so this page ranks by what fits your hardware rather than quoting scores. For a dated, citable snapshot of the whole landscape, see our Local AI Report #2: the best open-source coding models. If you want every local model scored and ranked for coding rather than sorted by tier, see the full benchmark ranking →.
The best coders you can download (workstation tier)
Open-weights state of the art. These need multi-GPU rigs or big unified memory all-in-VRAM — but note the MoE offload trick below the table.
| Model | VRAM (Q4) | Runs on | Context | License |
|---|---|---|---|---|
| Qwen3-Coder 480B-A35B (MoE) Open state of the art The largest dedicated open coder from Qwen, which positions it near Claude Sonnet 4 on agentic coding. Realistically an API or datacenter model. ollama pull qwen3-coder:480b-a35b | 290.6 GB | Multi-GPU server — or use an API Mac: 512 GB unified | 256K | Apache-2.0 |
| Devstral-2 123B Agentic specialist Mistral’s agentic flagship, built for multi-file edits and repo navigation. A dense 123B, so every parameter is read on every token: slower than the MoE options here. ollama pull devstral:123b | 75.1 GB | 2×48 GB GPUs / big unified memory Mac: 128 GB unified | 256K | Apache-2.0 |
| Qwen3-Coder-Next (80B-A3B MoE) Best self-hostable coder Released as Qwen3-Coder-Next: 80B total but only 3B active per token, on a hybrid-attention base whose long-context cache is cheap. All 80B must be resident, so plan for a 64 GB-class machine. ollama pull qwen3-coder:80b-a3b-q4 | 49.1 GB | 2×48 GB GPUs / big unified memory Mac: 96 GB unified | 256K | Apache-2.0 |
Best coding models for 16–24 GB GPUs
The consumer sweet spot — RTX 3090/4090, RX 7900 XTX, or a 32 GB Mac. This is where "self-hosted and genuinely good at code" starts.
| Model | VRAM (Q4) | Runs on | Context | License |
|---|---|---|---|---|
| Qwen 3.6 27B Best on a 24 GB card The strongest all-round coder that fits an RTX 3090/4090 comfortably at Q4 — the realistic daily driver for local agentic work. ollama pull qwen3.6:27b | 17.6 GB | 24 GB GPU (RTX 3090/4090) Mac: 24 GB unified | 256K | Apache-2.0 |
| Qwen3-Coder 30B-A3B (MoE) Fast MoE coder A dedicated coder with 3.3B active parameters, so it streams like a small model while holding 30B of knowledge. Fits a 24 GB card at Q4. ollama pull qwen3-coder:30b | 19.2 GB | 24 GB GPU (RTX 3090/4090) Mac: 32 GB unified | 256K | Apache-2.0 |
| Devstral Small 2 24B Agentic specialist Mistral’s small agentic coder, built to drive an editing loop. A dense 24B that fits a 24 GB card with room for context; on 16 GB it is a tight fit. ollama pull devstral-small-2:24b | 15.3 GB | 24 GB GPU (RTX 3090/4090) Mac: 24 GB unified | 256K | Apache-2.0 |
| Codestral 22B FIM autocomplete (license warning) Excellent fill-in-the-middle quality across 80+ languages — but the MNPL license forbids commercial/production use. Hobby projects only. ollama pull codestral:22b | 14.2 GB | 16 GB GPU (RTX 4060 Ti 16GB / 5060 Ti) Mac: 24 GB unified | 32K | MNPL-0.1 (non-production) |
Small coding models (4–12 GB) and autocomplete
Latency beats size for inline completions — a FIM-tuned small model next to your cursor outperforms a big chat model that takes two seconds to answer.
| Model | VRAM (Q4) | Runs on | Context | License |
|---|---|---|---|---|
| StarCoder 2 7B Best small autocomplete A fill-in-the-middle code model that fits any 8 GB GPU — the default pick for autocomplete on modest hardware. ollama pull starcoder2:7b | 5.1 GB | 8 GB GPU (RTX 3060/4060) Mac: 16 GB unified | 16K | BigCode OpenRAIL-M |
| StarCoder 2 15B Mid-size autocomplete A 600-language code model for 12 GB cards, for when you want more breadth than the 7B. ollama pull starcoder2:15b | 10.2 GB | 12 GB GPU (RTX 3060 12GB / 4070) Mac: 16 GB unified | 16K | BigCode OpenRAIL-M |
| StarCoder 2 3B Fastest autocomplete The lightweight option: small enough that ghost text keeps up with your typing. Pair it with a bigger chat model. ollama pull starcoder2:3b | 2.6 GB | 8 GB GPU (RTX 3060/4060) Mac: 16 GB unified | 16K | BigCode OpenRAIL-M |
The frontier you can’t download (and when to call it)
The largest open-weight coders — DeepSeek V4, GLM-5.x, and Kimi K2.x — want 80 GB-class multi-GPU hardware. For a solo developer they’re API models, not local models. The pattern that wins: a fast local coder for the inner loop (autocomplete, quick edits, private code) plus a frontier model over an API for the hardest problems. Our cloud AI directory lists where each frontier model is hosted and what it costs per token.
How to choose from these tables. Start from your VRAM, not from the leaderboard: pick the largest agentic coder your tier runs comfortably, add StarCoder 2 3B or StarCoder 2 7B for autocomplete if your card has headroom, and check the exact fit for your GPU — including quant options and context-length headroom — with the compatibility checker. Then wire it into an editor with our end-to-end setup guide.The hardware these tiers assume
The 24 GB tier is where local coding gets genuinely good, and two cards own it: the RTX 4090, and the used RTX 3090 as the budget route to the same VRAM. A 16 GB card such as the RTX 5060 Ti 16GB runs a mid-size coder or an autocomplete model plus a small chat model, but not Devstral Small 2 and an autocomplete model side by side. Street prices are well above launch MSRP right now, so compare before you buy.
No hardware? Rent the GPU first
Hardware falls short of the model you want? Rent a 24–48 GB GPU by the hour for a few dollars and test-drive the exact model before buying anything.
- RunPod — RTX 4090 ≈ $0.3–0.7/hr, A100/H100 by the hour; serverless per-second billing available
- Vast.ai — Marketplace of hosted GPUs — often the cheapest per hour (interruptible options)
- Lambda — On-demand A100/H100/B200 instances and clusters, per-hour billing
Affiliate links — we may earn a commission if you sign up, at no extra cost to you.
Full list on the cloud AI directory.
Frequently asked questions
What is the best local coding LLM for a 24 GB GPU (RTX 3090/4090)?
Qwen 3.6 27B is the best all-round coder that fits a 24 GB card at Q4 (~17.6 GB). Devstral Small 2 24B (~15.3 GB) is the agentic specialist, and Qwen3-Coder 30B-A3B (~19.2 GB) is the fast MoE option. Run one of them for agentic work and StarCoder 2 for autocomplete if you have the headroom.
What is the best coding model for 16 GB VRAM?
Devstral Small 2 24B needs ~15.3 GB at Q4 before the KV cache, so on a 16 GB card it fits only with a short context. For comfortable headroom, run StarCoder 2 15B (~10.2 GB) or StarCoder 2 7B (~5.1 GB) for autocomplete, and use the compatibility checker to test any agentic model against your exact card and context length.
Can I run Qwen3-Coder locally?
Yes, at three sizes. The 30B-A3B MoE (~19.2 GB at Q4) fits a 24 GB card. Qwen3-Coder-Next (80B-A3B) needs ~49.1 GB all-in-VRAM — a 64 GB-class machine — because every expert must be resident even though only 3B are active per token. The 480B-A35B flagship needs ~290.6 GB — datacenter or API territory.
Are local coding models as good as GitHub Copilot or Cursor?
For autocomplete and everyday edits on a 24 GB card, local models are close enough for many developers. For the hardest multi-file agentic tasks, frontier API models still lead. The winning setup is usually local-first with an API escape hatch.
Which local coding models allow commercial use?
Check the License column in each table, which comes from the model’s own record. Qwen3-Coder and Qwen 3.6 ship under Apache 2.0. Codestral 22B uses Mistral’s non-production license, so it is not for commercial work, and StarCoder2 uses BigCode OpenRAIL-M (commercial use with use-based restrictions).
Keep going
- Local AI Report #2 — the coding-model landscape, dated and citable
- Qwen3-Coder — full specs, quants, and variants
- Check what your GPU runs — free compatibility checker
More from this hub
- Run a Local Coding Agent End-to-End: Ollama + Qwen3-Coder + VS Code
- Local GitHub Copilot Alternatives: What Actually Replaces It Offline
- Evaluating a Local Coding Agent: Benchmarks Lie, Your Repo Doesn’t
- Tool Calling With Local Models: Making Agents Stop Breaking on Malformed JSON
- Local AI coding assistant — the full hub