Local AI Report #6 — Best Local Coding Models: Three Picks for Every Hardware Tier
The best local coding models for 8–16 GB, 24–32 GB and 128 GB machines: three picks per tier, what they run on, and how far behind Claude they score.
- Entry level (8–16 GB): K2 Horizon 7B. A 5.6 GB download that scores 21 on the independent Artificial Analysis index — nearly double Qwen3.5-9B's 11. Runners-up: Qwen3.5-9B (the easiest to install) and K2 Horizon 3.7B for 8 GB.
- Mid-range (24–32 GB): Qwen3.8-27B. Index 34, an 18 GB download, and about 21.9 GB at 64k context. Runners-up: K2 Horizon MoVA 36B-A4B (25, on a 32 GB card) and the dedicated coder Devstral Small 2.
- High-end (128 GB): Qwen3.8-Flash-Next. Index 40, a 120 GB download — tight on a DGX Spark or Strix Halo, too big for a 128 GB Mac. Runners-up: DeepSeek V4 Flash 0731 at 3-bit and Qwen3.8-27B at full precision.
- The gap to the subscriptions is real. The best model on one GPU scores 34; Claude Opus 5.5 scores 58 (index v4.3.2, checked 2026-10-08). That is the budget subscription tier, not the flagship.
- No model that fits 32 GB has an independent coding-agent score. Every coding benchmark quoted for the entry and mid-range picks is vendor-reported, and labelled that way.
- Context decides whether a pick is usable. Ollama starts you at 4k tokens below 24 GiB of VRAM; coding agents need at least 64,000.
The short answer: three picks per tier
Three memory tiers, three models each, ranked by the strongest evidence available: first the independent Artificial Analysis Intelligence Index, then the coding scores the model makers publish.
Three picks per hardware tier
Swipe sideways to see the whole chart →
| Tier | Pick | Download | Independent index |
|---|---|---|---|
| Entry level (8–16 GB) | 1. K2 Horizon 7B | 5.6 GB (Q4_K_M GGUF) | 21 |
| 2. Qwen3.5-9B | 6.6–7.6 GB | 11 | |
| 3. K2 Horizon 3.7B | 3.2 GB (Q4_K_M GGUF) | 16 | |
| Mid-range (24–32 GB) | 1. Qwen3.8-27B | 18 GB | 34 |
| 2. K2 Horizon MoVA 36B-A4B | 22.4 GB (Q4_K_M GGUF) | 25 | |
| 3. Devstral Small 2 24B | 15 GB | — | |
| High-end (128 GB) | 1. Qwen3.8-Flash-Next | 120 GB (Q4_K_M) / 105 GB (MLX) | 40 |
| 2. DeepSeek V4 Flash 0731 (3-bit) | 104 GB (UD-IQ3_XXS) | — | |
| 3. Qwen3.8-27B at Q8, full context | 30 GB | 34 |
Three kinds of figure appear on this page, and each is labelled where it appears:
- Independent — Artificial Analysis ran the test. The numbers to lead with.
- Vendor-reported — the model maker ran its own test on its own harness. Useful, never neutral.
- Community-reported — somebody else measured it on their hardware. We did not run these models for this report.
Not sure which tier you are in? The VRAM calculator checks your machine against any model.
Entry level: 8–16 GB
An 8 GB graphics card, a 12–16 GB card, or a laptop with 16 GB of unified memory. Fine for small, well-specified edits, single files and explanations. Thin for multi-file agent loops, because the context runs out before the model does.
1. K2 Horizon 7B — the best scored small model. From MBZUAI's Institute of Foundation Models, Apache 2.0. It scores 21 on the independent index, the highest of any model this size, against 11 for Qwen3.5-9B. Its makers report SWE-bench Verified 70.6 and Terminal-Bench 2.1 39.1, against 50.8 and 29.2 for Qwen3.5-9B in the same table. Treat a 7B model at 70 on SWE-bench Verified as a vendor claim until someone independent repeats it; the AA score is what says it is genuinely ahead. The Q4_K_M file is 5.6 GB; with its KV cache at 16k context it needs about 8.8 GB, so plan on a 12 GB card. The catch: there is no Ollama library tag yet. The official GGUFs need a llama.cpp build with K2 Horizon support — now merged into mainline llama.cpp — and an older Ollama or LM Studio may refuse to load them. 2. Qwen3.5-9B — the easy one. Index 11. One command:ollama pull qwen3.5:9b, a 6.6–7.6 GB download. If you want something that works this afternoon with no build steps, start here. It needs a 12 GB card or more at 16k context.
3. K2 Horizon 3.7B — for 8 GB. Index 16, a 3.2 GB file that needs about 6.4 GB at 16k context, and vendor-reported SWE-bench Verified 68.6 against 41.2 for Qwen3.5-4B — though on Terminal-Bench 2.1 the two are level (25.1 against 25.8). Same runtime caveat as the 7B. On Ollama alone, Qwen3.5-4B (qwen3.5:4b, 3.3–4.0 GB) is the fallback.
| Pick | Download | RTX 4060 (8 GB) | RTX 3060 (12GB) (12 GB) | RTX 5060 Ti 16GB (16 GB) |
|---|---|---|---|---|
| K2 Horizon 7B | 5.6 GB (Q4_K_M GGUF) | does not fit | fits | fits |
| Qwen3.5-9B | 6.6–7.6 GB | does not fit | fits | fits |
| K2 Horizon 3.7B | 3.2 GB (Q4_K_M GGUF) | fits | fits | fits |
- Granite 4.2 8B (AA 11) — 5.3 GB
- Qwen3.5-4B — 3.3–4.0 GB
- Gemma 4 12B — 7.7–8.0 GB
- gpt-oss-20b (16 GB only) — 14 GB
For laptops without a discrete GPU, and the Ollama default-tag trap that makes small Gemma models bigger than they look, see Report #3 on small models for 8 and 16 GB laptops. Per-size lists live on the 8 GB and 16 GB pages.
Mid-range: 24–32 GB
One 24 GB card (RTX 3090 or 4090) or a 32 GB RTX 5090. This is the sweet spot: the first tier where a coding agent can work in a real repository.
1. Qwen3.8-27B — the best model that fits one GPU. Index 34 at its default reasoning setting, 28 at medium, 26 at low. Apache 2.0, 262,144-token native context, image input. Alibaba reports SWE-bench Pro 61.7 and Terminal-Bench 2.1 73 on its model card — above Claude Opus 4.6 Max on the first (53.4), below it on the second (78.2).ollama pull qwen3.8:27b, 18 GB.
2. K2 Horizon MoVA 36B-A4B — fast, and a strong runner-up, on a 32 GB card. Index 25. A mixture-of-experts model with 36B parameters stored and 4B active per token, so it generates quickly. Its card reports Terminal-Bench 2.1 58.6 against 44.9 for Qwen3.6-35B-A3B and 51.7 for Muse Glimmer — and says those comparison columns are Artificial Analysis's runs. The Q4_K_M GGUF is 22.4 GB, and its plain attention cache is not small: at 32k context it needs about 29.6 GB, which fits, tight on an RTX 5090 and does not fit on a 24 GB card. Same runtime caveat as the small K2 models: llama.cpp, no Ollama tag yet.
3. Devstral Small 2 24B — the dedicated coder. Mistral's 24B coding model, Apache 2.0, built for agent work in its Mistral Vibe CLI. AA does not score it; Mistral reports SWE-bench Verified 68 and Terminal-Bench 2 22.5 (card). At 15 GB it leaves the most room on a 24 GB card for long context. ollama pull devstral-small-2:24b.
| Pick | Download | RTX 3090 (24 GB) | RTX 4090 (24 GB) | RTX 5090 (32 GB) |
|---|---|---|---|---|
| Qwen3.8-27B | 18 GB | fits | fits | fits |
| K2 Horizon MoVA 36B-A4B | 22.4 GB (Q4_K_M GGUF) | does not fit | does not fit | fits, tight |
| Devstral Small 2 24B | 15 GB | fits | fits | fits |
- Qwen3.6-35B-A3B (AA 18; fast MoE, 32 GB card) — 23–24 GB
- Gemma 4 31B (AA 15) — 19–20 GB
- Muse Glimmer 30B (AA 17) — —
- Everything in the entry tier, at higher precision
Three settings that decide how Qwen3.8-27B feels
1. Reasoning effort. The default is xhigh, and it overthinks: AA notes it used 200 million output tokens on its index against a median of 82 million, and Simon Willison waited 21 minutes for one prompt that took 137 seconds with reasoning off (his write-up). Start at medium. Qwen's card warns that in agent loops lower effort "does not always reduce overall task completion time", so time whole tasks.
2. Context. At least 64,000 tokens. That fits a 24 GB card; 128k needs 32 GB or an 8-bit KV cache.
3. Speculative decoding. Ollama publishes -mtp builds (qwen3.8:27b-mtp-q4_K_M) that use the model's built-in prediction head. Tuned community stacks run 2× to 3.2× faster than stock.
High-end: 128 GB
A DGX Spark, an AMD Strix Halo mini PC or a 128 GB Mac. More memory buys bigger models and longer context — but not, today, a clearly better coder that runs comfortably.
1. Qwen3.8-Flash-Next — the best scored model in reach. Index 40, 6 points above the 27B. 125B parameters with 6B active, plus a 51B n-gram embedding table, under the Qwen Community License. Ollama lists a 120 GB Q4_K_M build and a 105 GB MLX build. Our engine calls the Q4_K_M build fits, tight on a DGX Spark at 32k context, and does not fit on a 128 GB Mac, where macOS leaves about 96 GB for the GPU by default. 2. DeepSeek V4 Flash 0731, 3-bit — the only one with an independent coding-agent score. MIT licence. AA's Coding Agent Index scores the hosted model at 39, level with Claude Haiku 5.5. A 3-bit local build will score lower, and nobody has published by how much. The 3-bit files are 104.2 GB (UD-IQ3_XXS) and 128.2 GB (UD-Q3_K_XL); take the smaller one. 3. Qwen3.8-27B at Q8, full context — the comfortable choice. The mid-range winner (index 34) at 8-bit, which costs less quality than 4-bit. With the full 256k-token context it needs about 47.5 GB, so a 128 GB machine can run it with long context and several sessions at once. Of the three, it is the one that is not tight.| Pick | Download | DGX Spark (128 GB) | Ryzen AI Max+ 395 (128 GB) | Apple M5 Max (128 GB) |
|---|---|---|---|---|
| Qwen3.8-Flash-Next | 120 GB (Q4_K_M) / 105 GB (MLX) | fits, tight | fits, tight | does not fit |
| DeepSeek V4 Flash 0731 (3-bit) | 104 GB (UD-IQ3_XXS) | fits | fits | does not fit |
| Qwen3.8-27B at Q8, full context | 30 GB | fits | fits | fits |
- Qwen3-Coder-Next 80B-A3B (coding specialist, AA 9) — 52 GB
- Qwen3.5 122B-A10B (AA 16) — 81 GB
- gpt-oss-120b (AA 12) — 65 GB
- Every mid-range pick, several at once
Beyond 128 GB, the open models that lead the independent charts — GLM-5.3, Kimi K3, MiMo-V2.6-Pro, scoring 44 to 46 — do not fit home hardware. Kimi K3's weights alone are 1.56 TB. To try them, rent a GPU by the hour.
Need more VRAM than you own? Spin up a cloud GPU big enough for any model in minutes — pay for the hours you use instead of buying a card.
Affiliate links — we may earn a commission if you sign up, at no extra cost to you.
Marketplace pricing — the cheapest per-hour rates for spot/interruptible GPUs.
Rent GPUs on Vast.ai →On-demand pods with a simple UI — good for a quick one-off inference or fine-tune.
Rent GPUs on RunPod →How far behind the subscriptions are they?
The best model that fits one GPU scores 34. Claude Opus 5.5 scores 58. That 24-point gap is the honest starting point for anyone thinking about cancelling a subscription.
How far behind: independent scores
Swipe sideways to see the whole chart →
A score of 34 puts the mid-range winner level with the budget tier of the subscription families: Claude Haiku 5.5 at medium effort scores 34, GPT-6 Luna at high effort 33. The flagships — Opus 5.5 at 58, Sonnet 5.5 at 56 — are a different class.
The coding-specific view has a hole in it. AA's Coding Agent Index (v1.5) runs each model inside a real agent harness on three coding benchmarks. On 2026-10-08 it listed 37 entries. None of them is a model that fits in 32 GB. The top entry, Claude Code with Sonnet 5.5, scores 68; the best open model, GLM-5.3, 54. The one entry in reach of a 128 GB box, DeepSeek V4 Flash 0731 at 39, is 29 points under the top. Discount every vendor number. Anthropic reports 70.6 for Sonnet 5.5 on Terminal-Bench 4.0; AA's own run measured 66.2. For Opus 5.5 it is 66.4 against 63.1. That is 3.3 to 4.4 points from a vendor with a good reputation for reporting. The harness matters too: Qwen3.8-Flash-Next's card reports DeepSWE as "the highest score across the two harnesses" it tried. And passing tests is not shipping. METR had 4 maintainers of 3 SWE-bench Verified repositories review 296 AI-written patches that passed the benchmark's tests. Roughly half would not have been merged (METR's note). Cost and limits. Local removes the per-token bill and the session and weekly allowances; Anthropic's plans state that usage limits apply. It does not remove the limits of your memory (context), your bandwidth (speed) or the model (quality). For comparison, Claude Pro is $20 a month, Claude Max from $100 a month and Cursor Pro $20 a month (vendor pages, checked 2026-10-08). The cost calculator works out the break-even with your own numbers.Context and speed: what the tiers really buy
Memory decides what runs; context decides whether it is useful. Qwen3.8-27B's weights are about 16.8 GB at Q4_K_M, but coding agents do not run at short context.
Context length is the memory bill
Swipe sideways to see the whole chart →
Those totals come from this site's engine, using the model's published configuration: only 16 of its 64 layers keep a KV cache, 64 KB per token in F16. Hardware Corner measured the same model at 18, 22, 26 and 34 GB. On the site's fit rule: RTX 3090 fits, tight at 64k and does not fit at 128k; RTX 4090 fits, tight at 64k and does not fit at 128k; RTX 5090 fits at 64k and fits at 128k.
Speed is a bandwidth story
Swipe sideways to see the whole chart →
For a dense model each token reads every weight once, so speed is capped at memory bandwidth divided by the weights: about 55.8 tokens per second on an RTX 3090 and 106.8 on an RTX 5090, by our engine. Stock llama.cpp lands at 70 to 77 per cent of the ceiling on the machines Hardware Corner measured. Speculative decoding bends the rule: the syv-ai stack reached 127 tokens per second on one RTX 3090 and about 1,035 across 64 requests (repo). The RTX 5090, DGX Spark and Strix Halo (24–36 tuned, 8–11 stock) tuned figures come from one aggregator, Context Studios. We did not run any of these.
That is also why the 128 GB tier feels slower than its size suggests: unified-memory boxes have a fraction of a graphics card's bandwidth. Mixture-of-experts picks — K2 Horizon MoVA, Qwen3.8-Flash-Next — read only their active parameters per token, which is how they stay usable.
Going fully offline
The fully offline coding stack
Swipe sideways to see the whole chart →
After the one-time download, nothing needs the network. That rules out web search and fetch tools, package installs without a local mirror, and anything that signs in to a service.
# 1. One-time, online: install Ollama and pull a model from your tier.
ollama pull qwen3.8:27b
# 2. Serve with a context long enough for a coding agent.
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
# 3. Point Claude Code at the local server, and switch off its non-essential traffic.
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
claude --model qwen3.8:27b # or: ollama launch claude
The variables are the ones in Ollama's Claude Code guide, checked 2026-10-08. Run ollama ps and check that the CONTEXT column shows what you asked for: Ollama defaults to 4k below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k above. Ollama's Anthropic-compatible API does not support prompt caching, so long sessions re-read their context; expect slower prompt processing than a hosted run. Qwen Code, OpenCode, Cline, Aider and Continue all work against the same server — see the coding-agents guide. Then unplug the network cable and run a real task: anything that quietly needed the internet shows up in ten minutes. More in offline AI.
Honest limits and methodology
- We did not run these models. Speeds are Hardware Corner's (stock), the syv-ai repository's (tuned RTX 3090) and one aggregator's (other tuned rows). Memory totals, fit verdicts and bandwidth ceilings are this site's engine; Hardware Corner's measurements agree with it.
- The independent index is general, not coding-only. Dedicated coders like Devstral Small 2 and Qwen3-Coder-Next score low on it or are not scored, and can still be good at code. That is why vendor coding scores sit beside it, labelled.
- K2 Horizon is new and lightly tested. Strong independent scores, but no Ollama tag yet and fresh llama.cpp support. If you need it working today with no build steps, take each tier's Ollama pick.
- Quantisation costs quality. Every score here is for the full-precision model; nobody has benchmarked the 4-bit builds you will run, and the 3-bit DeepSeek is a particular worry.
- Benchmarks are not your repository. Run your last three real tasks through a local setup and count how many you would have accepted.
Frequently Asked Questions
What is the best local coding model right now?
It depends on your memory. On 8–16 GB, K2 Horizon 7B (Artificial Analysis index 21). On a 24–32 GB GPU, Qwen3.8-27B (34). On a 128 GB machine, Qwen3.8-Flash-Next (40), though it is a tight fit. Qwen3.8-27B is the best overall balance of quality and ease of running.
What is the best coding model for an 8 GB GPU or laptop?
K2 Horizon 3.7B, a 3.2 GB file that scores 16 on the independent index. If you want an Ollama one-liner instead, Qwen3.5-4B is a 3.3–4.0 GB download. Use either for small, well-specified edits rather than multi-file agent work.
What is the best coding model for a 12 or 16 GB GPU?
K2 Horizon 7B: a 5.6 GB file scoring 21 on the independent index, nearly double Qwen3.5-9B's 11. It needs a recent llama.cpp build because Ollama has no tag for it yet. The easiest alternative is Qwen3.5-9B, installed with ollama pull qwen3.5:9b.
What is the best coding model for an RTX 3090, 4090 or 5090?
Qwen3.8-27B. It scores 34 on the independent index, the highest of any model that fits one consumer GPU, and needs about 21.9 GB at 64k context. K2 Horizon MoVA 36B-A4B is a faster runner-up on a 32 GB RTX 5090, and Devstral Small 2 leaves the most room for long context on a 24 GB card.
What is the best coding model for a 128 GB DGX Spark, Strix Halo or Mac?
Qwen3.8-Flash-Next scores 40, the best of anything in reach, but its 120 GB build is tight on a DGX Spark or Strix Halo and too big for a 128 GB Mac at default settings. Qwen3.8-27B at 8-bit with full context is the comfortable choice on all three.
How much VRAM do I need for Qwen3.8-27B?
About 17.8 GB at short context, 21.9 GB at 64k and 26.2 GB at 128k, at Q4_K_M with a 16-bit KV cache, by this site's engine; Hardware Corner measured 18, 22 and 26 GB. A 24 GB GPU handles 64k and a 32 GB GPU handles 128k.
How good are local coding models compared with Claude, GPT and Gemini?
On the independent Artificial Analysis index, the best model that fits one GPU scores 34 and the top subscription models score 53 to 58. That matches the budget tier, such as Claude Haiku 5.5 at medium effort, not the flagships.
Can I use Claude Code with a local model, fully offline?
Yes. Pull the model once with Ollama, then set ANTHROPIC_BASE_URL to http://localhost:11434, ANTHROPIC_AUTH_TOKEN to ollama and ANTHROPIC_API_KEY to an empty string. Set the context to at least 64,000 tokens. Web search and fetch tools will not work offline.
Why does my local coding agent forget the task halfway through?
Usually context. Ollama defaults to 4k tokens below 24 GiB of VRAM, while coding tools need at least 64,000. Set OLLAMA_CONTEXT_LENGTH=64000, restart the server and confirm with ollama ps.
Will a local model replace my coding subscription?
For some work, yes; for hard problems, not yet. The independent gap between the best home model and the best subscription model is 24 points. Most people do best with a hybrid: local for private and bulk work, a subscription for the problems that stall.
Run it yourself: VRAM calculator · Local vs cloud cost · Offline AI · Coding agents guide
Earlier reports: Best open-source coding models (July) · Small LLMs for 8GB and 16GB laptops · Best laptops for local AI · Europe's AI compute gap
How the numbers are made: Methodology