Report #6 · published October 8, 2026

Local AI Report #6 — Best Local Coding Models: Three Picks for Every Hardware Tier

The best local coding models for 8–16 GB, 24–32 GB and 128 GB machines: three picks per tier, what they run on, and how far behind Claude they score.

Jakub Rusinowski

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
TL;DR
  • Entry level (8–16 GB): K2 Horizon 7B. A 5.6 GB download that scores 21 on the independent Artificial Analysis index — nearly double Qwen3.5-9B's 11. Runners-up: Qwen3.5-9B (the easiest to install) and K2 Horizon 3.7B for 8 GB.
  • Mid-range (24–32 GB): Qwen3.8-27B. Index 34, an 18 GB download, and about 21.9 GB at 64k context. Runners-up: K2 Horizon MoVA 36B-A4B (25, on a 32 GB card) and the dedicated coder Devstral Small 2.
  • High-end (128 GB): Qwen3.8-Flash-Next. Index 40, a 120 GB download — tight on a DGX Spark or Strix Halo, too big for a 128 GB Mac. Runners-up: DeepSeek V4 Flash 0731 at 3-bit and Qwen3.8-27B at full precision.
  • The gap to the subscriptions is real. The best model on one GPU scores 34; Claude Opus 5.5 scores 58 (index v4.3.2, checked 2026-10-08). That is the budget subscription tier, not the flagship.
  • No model that fits 32 GB has an independent coding-agent score. Every coding benchmark quoted for the entry and mid-range picks is vendor-reported, and labelled that way.
  • Context decides whether a pick is usable. Ollama starts you at 4k tokens below 24 GiB of VRAM; coding agents need at least 64,000.

The short answer: three picks per tier

Three memory tiers, three models each, ranked by the strongest evidence available: first the independent Artificial Analysis Intelligence Index, then the coding scores the model makers publish.

Three picks per hardware tier

Best local coding models for each hardware tierThree columns, one per hardware tier. Entry level, 8–16 GB: 1. K2 Horizon 7B, 5.6 GB (Q4_K_M GGUF), index 21; 2. Qwen3.5-9B, 6.6–7.6 GB, index 11; 3. K2 Horizon 3.7B, 3.2 GB (Q4_K_M GGUF), index 16. Example machines: NVIDIA GeForce RTX 4060, NVIDIA GeForce RTX 3060 (12GB), NVIDIA GeForce RTX 5060 Ti 16GB. Mid-range, 24–32 GB: 1. Qwen3.8-27B, 18 GB, index 34; 2. K2 Horizon MoVA 36B-A4B, 22.4 GB (Q4_K_M GGUF), index 25; 3. Devstral Small 2 24B, 15 GB, not scored. Example machines: NVIDIA GeForce RTX 3090, NVIDIA GeForce RTX 4090, NVIDIA GeForce RTX 5090. High-end, 128 GB: 1. Qwen3.8-Flash-Next, 120 GB (Q4_K_M) / 105 GB (MLX), index 40; 2. DeepSeek V4 Flash 0731 (3-bit), 104 GB (UD-IQ3_XXS), not scored; 3. Qwen3.8-27B at Q8, full context, 30 GB, index 34. Example machines: NVIDIA DGX Spark, AMD Ryzen AI Max+ 395, Apple M5 Max.The best local coding models, three per hardware tierBadge: Artificial Analysis Intelligence Index (independent). n/s = not scored by AA.Entry level8–16 GB1. K2 Horizon 7B5.6 GB (Q4_K_M GGUF)212. Qwen3.5-9B6.6–7.6 GB113. K2 Horizon 3.7B3.2 GB (Q4_K_M GGUF)16RTX 4060 · RTX 3060 (12GB) · RTX5060 Ti 16GBMid-range24–32 GB1. Qwen3.8-27B18 GB342. K2 Horizon MoVA 36B-A4B22.4 GB (Q4_K_M GGUF)253. Devstral Small 2 24B15 GBn/sRTX 3090 · RTX 4090 · RTX 5090High-end128 GB1. Qwen3.8-Flash-Next120 GB (Q4_K_M) / 105 GB (MLX)402. DeepSeek V4 Flash 0731(3-bit)104 GB (UD-IQ3_XXS)n/s3. Qwen3.8-27B at Q8, fullcontext30 GB34DGX Spark · Ryzen AI Max+ 395 ·Apple M5 Max

Swipe sideways to see the whole chart →

Sources: Artificial Analysis Intelligence Index (independent scores); Ollama library and official GGUF listings (download sizes). Checked 2026-10-08.
In words: K2 Horizon 7B leads at 8–16 GB (index 21), Qwen3.8-27B at 24–32 GB (34), and Qwen3.8-Flash-Next at 128 GB (40).
TierPickDownloadIndependent index
Entry level (8–16 GB)1. K2 Horizon 7B5.6 GB (Q4_K_M GGUF)21
2. Qwen3.5-9B6.6–7.6 GB11
3. K2 Horizon 3.7B3.2 GB (Q4_K_M GGUF)16
Mid-range (24–32 GB)1. Qwen3.8-27B18 GB34
2. K2 Horizon MoVA 36B-A4B22.4 GB (Q4_K_M GGUF)25
3. Devstral Small 2 24B15 GB—
High-end (128 GB)1. Qwen3.8-Flash-Next120 GB (Q4_K_M) / 105 GB (MLX)40
2. DeepSeek V4 Flash 0731 (3-bit)104 GB (UD-IQ3_XXS)—
3. Qwen3.8-27B at Q8, full context30 GB34
Independent index: Artificial Analysis Intelligence Index v4.3.2, checked 2026-10-08 — a general agentic index, not a pure coding test. "—" means AA does not score the model. Downloads are Ollama library sizes or official GGUF files.

Three kinds of figure appear on this page, and each is labelled where it appears:

  • Independent — Artificial Analysis ran the test. The numbers to lead with.
  • Vendor-reported — the model maker ran its own test on its own harness. Useful, never neutral.
  • Community-reported — somebody else measured it on their hardware. We did not run these models for this report.

Not sure which tier you are in? The VRAM calculator checks your machine against any model.

Entry level: 8–16 GB

An 8 GB graphics card, a 12–16 GB card, or a laptop with 16 GB of unified memory. Fine for small, well-specified edits, single files and explanations. Thin for multi-file agent loops, because the context runs out before the model does.

1. K2 Horizon 7B — the best scored small model. From MBZUAI's Institute of Foundation Models, Apache 2.0. It scores 21 on the independent index, the highest of any model this size, against 11 for Qwen3.5-9B. Its makers report SWE-bench Verified 70.6 and Terminal-Bench 2.1 39.1, against 50.8 and 29.2 for Qwen3.5-9B in the same table. Treat a 7B model at 70 on SWE-bench Verified as a vendor claim until someone independent repeats it; the AA score is what says it is genuinely ahead. The Q4_K_M file is 5.6 GB; with its KV cache at 16k context it needs about 8.8 GB, so plan on a 12 GB card. The catch: there is no Ollama library tag yet. The official GGUFs need a llama.cpp build with K2 Horizon support — now merged into mainline llama.cpp — and an older Ollama or LM Studio may refuse to load them. 2. Qwen3.5-9B — the easy one. Index 11. One command: ollama pull qwen3.5:9b, a 6.6–7.6 GB download. If you want something that works this afternoon with no build steps, start here. It needs a 12 GB card or more at 16k context. 3. K2 Horizon 3.7B — for 8 GB. Index 16, a 3.2 GB file that needs about 6.4 GB at 16k context, and vendor-reported SWE-bench Verified 68.6 against 41.2 for Qwen3.5-4B — though on Terminal-Bench 2.1 the two are level (25.1 against 25.8). Same runtime caveat as the 7B. On Ollama alone, Qwen3.5-4B (qwen3.5:4b, 3.3–4.0 GB) is the fallback.
PickDownloadRTX 4060 (8 GB)RTX 3060 (12GB) (12 GB)RTX 5060 Ti 16GB (16 GB)
K2 Horizon 7B5.6 GB (Q4_K_M GGUF)does not fitfitsfits
Qwen3.5-9B6.6–7.6 GBdoes not fitfitsfits
K2 Horizon 3.7B3.2 GB (Q4_K_M GGUF)fitsfitsfits
Verdicts from this site's engine at 16k context, counting weights, KV cache and runtime overhead against each machine's usable memory, checked 2026-10-08. Also runs on this tier:
  • Granite 4.2 8B (AA 11) — 5.3 GB
  • Qwen3.5-4B — 3.3–4.0 GB
  • Gemma 4 12B — 7.7–8.0 GB
  • gpt-oss-20b (16 GB only) — 14 GB

For laptops without a discrete GPU, and the Ollama default-tag trap that makes small Gemma models bigger than they look, see Report #3 on small models for 8 and 16 GB laptops. Per-size lists live on the 8 GB and 16 GB pages.

Mid-range: 24–32 GB

One 24 GB card (RTX 3090 or 4090) or a 32 GB RTX 5090. This is the sweet spot: the first tier where a coding agent can work in a real repository.

1. Qwen3.8-27B — the best model that fits one GPU. Index 34 at its default reasoning setting, 28 at medium, 26 at low. Apache 2.0, 262,144-token native context, image input. Alibaba reports SWE-bench Pro 61.7 and Terminal-Bench 2.1 73 on its model card — above Claude Opus 4.6 Max on the first (53.4), below it on the second (78.2). ollama pull qwen3.8:27b, 18 GB. 2. K2 Horizon MoVA 36B-A4B — fast, and a strong runner-up, on a 32 GB card. Index 25. A mixture-of-experts model with 36B parameters stored and 4B active per token, so it generates quickly. Its card reports Terminal-Bench 2.1 58.6 against 44.9 for Qwen3.6-35B-A3B and 51.7 for Muse Glimmer — and says those comparison columns are Artificial Analysis's runs. The Q4_K_M GGUF is 22.4 GB, and its plain attention cache is not small: at 32k context it needs about 29.6 GB, which fits, tight on an RTX 5090 and does not fit on a 24 GB card. Same runtime caveat as the small K2 models: llama.cpp, no Ollama tag yet. 3. Devstral Small 2 24B — the dedicated coder. Mistral's 24B coding model, Apache 2.0, built for agent work in its Mistral Vibe CLI. AA does not score it; Mistral reports SWE-bench Verified 68 and Terminal-Bench 2 22.5 (card). At 15 GB it leaves the most room on a 24 GB card for long context. ollama pull devstral-small-2:24b.
PickDownloadRTX 3090 (24 GB)RTX 4090 (24 GB)RTX 5090 (32 GB)
Qwen3.8-27B18 GBfitsfitsfits
K2 Horizon MoVA 36B-A4B22.4 GB (Q4_K_M GGUF)does not fitdoes not fitfits, tight
Devstral Small 2 24B15 GBfitsfitsfits
Verdicts from this site's engine at 32k context, counting weights, KV cache and runtime overhead against each machine's usable memory, checked 2026-10-08. Also runs on this tier:
  • Qwen3.6-35B-A3B (AA 18; fast MoE, 32 GB card) — 23–24 GB
  • Gemma 4 31B (AA 15) — 19–20 GB
  • Muse Glimmer 30B (AA 17) — —
  • Everything in the entry tier, at higher precision

Three settings that decide how Qwen3.8-27B feels

1. Reasoning effort. The default is xhigh, and it overthinks: AA notes it used 200 million output tokens on its index against a median of 82 million, and Simon Willison waited 21 minutes for one prompt that took 137 seconds with reasoning off (his write-up). Start at medium. Qwen's card warns that in agent loops lower effort "does not always reduce overall task completion time", so time whole tasks. 2. Context. At least 64,000 tokens. That fits a 24 GB card; 128k needs 32 GB or an 8-bit KV cache. 3. Speculative decoding. Ollama publishes -mtp builds (qwen3.8:27b-mtp-q4_K_M) that use the model's built-in prediction head. Tuned community stacks run 2× to 3.2× faster than stock.

NVIDIA GeForce RTX 5090 32GB
32 GB VRAM · 575 W board power
2026 prices are volatile — check the current listing.

High-end: 128 GB

A DGX Spark, an AMD Strix Halo mini PC or a 128 GB Mac. More memory buys bigger models and longer context — but not, today, a clearly better coder that runs comfortably.

1. Qwen3.8-Flash-Next — the best scored model in reach. Index 40, 6 points above the 27B. 125B parameters with 6B active, plus a 51B n-gram embedding table, under the Qwen Community License. Ollama lists a 120 GB Q4_K_M build and a 105 GB MLX build. Our engine calls the Q4_K_M build fits, tight on a DGX Spark at 32k context, and does not fit on a 128 GB Mac, where macOS leaves about 96 GB for the GPU by default. 2. DeepSeek V4 Flash 0731, 3-bit — the only one with an independent coding-agent score. MIT licence. AA's Coding Agent Index scores the hosted model at 39, level with Claude Haiku 5.5. A 3-bit local build will score lower, and nobody has published by how much. The 3-bit files are 104.2 GB (UD-IQ3_XXS) and 128.2 GB (UD-Q3_K_XL); take the smaller one. 3. Qwen3.8-27B at Q8, full context — the comfortable choice. The mid-range winner (index 34) at 8-bit, which costs less quality than 4-bit. With the full 256k-token context it needs about 47.5 GB, so a 128 GB machine can run it with long context and several sessions at once. Of the three, it is the one that is not tight.
PickDownloadDGX Spark (128 GB)Ryzen AI Max+ 395 (128 GB)Apple M5 Max (128 GB)
Qwen3.8-Flash-Next120 GB (Q4_K_M) / 105 GB (MLX)fits, tightfits, tightdoes not fit
DeepSeek V4 Flash 0731 (3-bit)104 GB (UD-IQ3_XXS)fitsfitsdoes not fit
Qwen3.8-27B at Q8, full context30 GBfitsfitsfits
Verdicts from this site's engine at 32k context, counting weights, KV cache and runtime overhead against each machine's usable memory, checked 2026-10-08. Also runs on this tier:
  • Qwen3-Coder-Next 80B-A3B (coding specialist, AA 9) — 52 GB
  • Qwen3.5 122B-A10B (AA 16) — 81 GB
  • gpt-oss-120b (AA 12) — 65 GB
  • Every mid-range pick, several at once

Beyond 128 GB, the open models that lead the independent charts — GLM-5.3, Kimi K3, MiMo-V2.6-Pro, scoring 44 to 46 — do not fit home hardware. Kimi K3's weights alone are 1.56 TB. To try them, rent a GPU by the hour.

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
128 GB VRAM · 120 W board power
2026 prices are volatile — check the current listing.

Rent a GPU by the hour

Need more VRAM than you own? Spin up a cloud GPU big enough for any model in minutes — pay for the hours you use instead of buying a card.

Affiliate links — we may earn a commission if you sign up, at no extra cost to you.

Vast.ai

Marketplace pricing — the cheapest per-hour rates for spot/interruptible GPUs.

Rent GPUs on Vast.ai →
RunPod

On-demand pods with a simple UI — good for a quick one-off inference or fine-tune.

Rent GPUs on RunPod →

How far behind the subscriptions are they?

The best model that fits one GPU scores 34. Claude Opus 5.5 scores 58. That 24-point gap is the honest starting point for anyone thinking about cancelling a subscription.

How far behind: independent scores

Independent index scores: subscription models versus models that run at homeBar chart of the Artificial Analysis Intelligence Index v4.3.2. Claude Opus 5.5 (max) 58; Claude Sonnet 5.5 (max) 56; GPT-6 Astra (max) 53; Gemini 4 Argon (high) 53; MiMo-V2.6-Pro 46; GLM-5.3 (max) 45; Kimi K3 (max) 44; Qwen3.8 27B (xhigh) 34; K2 Horizon MoVA 36B A4B 25; Qwen3.6 35B A3B 18; Muse Glimmer (high) 17; Gemma 4 31B 15; K2 Horizon 7B 21; K2 Horizon 3.7B 16; Qwen3.5 9B 11; Granite 4.2 8B 11. The best model that fits one consumer GPU, Qwen3.8 27B, scores 34 at its highest reasoning setting, 28 at medium and 26 at low, against 58 for the top subscription model.Independent score: subscription models vs. what fits at homeArtificial Analysis Intelligence Index v4.3.2 — a general agentic index, not a coding test.0102030405060SUBSCRIPTION MODELS (CLOSED WEIGHTS)Claude Opus 5.5 (max)58Claude Sonnet 5.5 (max)56GPT-6 Astra (max)53Gemini 4 Argon (high)53OPEN WEIGHTS, SERVER-CLASSMiMo-V2.6-Pro46GLM-5.3 (max)45Kimi K3 (max)44FITS A 24–32 GB MACHINE AT 4-BITQwen3.8 27B (xhigh)3428 at medium · 26 at low (ticks)K2 Horizon MoVA 36B A4B25Qwen3.6 35B A3B18Muse Glimmer (high)17Gemma 4 31B15FITS AN 8–16 GB MACHINE AT 4-BITK2 Horizon 7B21K2 Horizon 3.7B16Qwen3.5 9B11Granite 4.2 8B11best that fits one GPU: 34 · top subscription: 58

Swipe sideways to see the whole chart →

Source: Artificial Analysis Intelligence Index v4.3.2 — independent, and a general agentic index rather than a coding test. Grouping by memory fit is ours. Checked 2026-10-08.
In words: the subscription flagships score 53 to 58, the open server-class models 44 to 46, the best mid-range pick 34 and the best entry pick 21.

A score of 34 puts the mid-range winner level with the budget tier of the subscription families: Claude Haiku 5.5 at medium effort scores 34, GPT-6 Luna at high effort 33. The flagships — Opus 5.5 at 58, Sonnet 5.5 at 56 — are a different class.

The coding-specific view has a hole in it. AA's Coding Agent Index (v1.5) runs each model inside a real agent harness on three coding benchmarks. On 2026-10-08 it listed 37 entries. None of them is a model that fits in 32 GB. The top entry, Claude Code with Sonnet 5.5, scores 68; the best open model, GLM-5.3, 54. The one entry in reach of a 128 GB box, DeepSeek V4 Flash 0731 at 39, is 29 points under the top. Discount every vendor number. Anthropic reports 70.6 for Sonnet 5.5 on Terminal-Bench 4.0; AA's own run measured 66.2. For Opus 5.5 it is 66.4 against 63.1. That is 3.3 to 4.4 points from a vendor with a good reputation for reporting. The harness matters too: Qwen3.8-Flash-Next's card reports DeepSWE as "the highest score across the two harnesses" it tried. And passing tests is not shipping. METR had 4 maintainers of 3 SWE-bench Verified repositories review 296 AI-written patches that passed the benchmark's tests. Roughly half would not have been merged (METR's note). Cost and limits. Local removes the per-token bill and the session and weekly allowances; Anthropic's plans state that usage limits apply. It does not remove the limits of your memory (context), your bandwidth (speed) or the model (quality). For comparison, Claude Pro is $20 a month, Claude Max from $100 a month and Cursor Pro $20 a month (vendor pages, checked 2026-10-08). The cost calculator works out the break-even with your own numbers.

Context and speed: what the tiers really buy

Memory decides what runs; context decides whether it is useful. Qwen3.8-27B's weights are about 16.8 GB at Q4_K_M, but coding agents do not run at short context.

Context length is the memory bill

Memory needed by Qwen3.8-27B as context growsStacked bars of memory needed by Qwen3.8-27B at Q4_K_M. 4k context: 17.8 GB with an F16 KV cache, 17.7 GB with an 8-bit KV cache; 64k context: 21.9 GB with an F16 KV cache, 19.7 GB with an 8-bit KV cache; 128k context: 26.2 GB with an F16 KV cache, 21.9 GB with an 8-bit KV cache; 256k context: 34.8 GB with an F16 KV cache, 26.2 GB with an 8-bit KV cache. Dashed lines mark 16, 24 and 32 GB.Context length is the memory bill — Qwen3.8-27B at Q4_K_MGB needed, text-only. Each pair: F16 KV cache (left) vs. an 8-bit KV cache (right).Weights, Q4_K_M (16.8 GB)KV cache (64 KB per token in F16)Runtime overhead0816243217.817.74k contextF16 KV8-bit KV21.919.764k contextF16 KV8-bit KV26.221.9128k contextF16 KV8-bit KV34.826.2256k contextF16 KV8-bit KV16 GB24 GB card32 GB card

Swipe sideways to see the whole chart →

Source: computed by this site's engine from Qwen3.8-27B's published config (text-only; the optional vision projector is extra). Hardware Corner measured 18 / 22 / 26 / 34 GB on the same model. Checked 2026-10-08.
In words: Qwen3.8-27B needs about 17.8 GB at 4k context, 21.9 GB at 64k, 26.2 GB at 128k and 34.8 GB at 256k; an 8-bit KV cache brings 128k down to 21.9 GB.

Those totals come from this site's engine, using the model's published configuration: only 16 of its 64 layers keep a KV cache, 64 KB per token in F16. Hardware Corner measured the same model at 18, 22, 26 and 34 GB. On the site's fit rule: RTX 3090 fits, tight at 64k and does not fit at 128k; RTX 4090 fits, tight at 64k and does not fit at 128k; RTX 5090 fits at 64k and fits at 128k.

Speed is a bandwidth story

Qwen3.8-27B speed by hardware: bandwidth ceiling, stock and tunedQwen3.8-27B decode speed by hardware. RTX 3090: ceiling 55.8 tokens per second, stock 40.3, tuned 127. RTX 4090: ceiling 60.1 tokens per second, stock 46.2, no tuned figure. RTX 5090: ceiling 106.8 tokens per second, stock 74.8, tuned 148. DGX Spark (GB10): ceiling 16.3 tokens per second, stock 12.2, tuned 33 to 35. Strix Halo (128 GB): ceiling 15.3 tokens per second, stock 8 to 11, tuned 24 to 36.Qwen3.8-27B tokens per second: ceiling, stock and tunedStock lands at 70–77% of the ceiling; speculative decoding beats it.Bandwidth ceiling (engine)Stock llama.cppCommunity-tuned (MTP / DFlash)04080120160RTX 309024 GB · 936 GB/sceiling 55.840.3 stock127 tunedRTX 409024 GB · 1008 GB/sceiling 60.146.2 stocktuned: no figure in our sourcesRTX 509032 GB · 1792 GB/sceiling 106.874.8 stock148 tunedDGX Spark (GB10)128 GB · 273 GB/sceiling 16.312.2 stock33–35 tunedStrix Halo (128 GB)128 GB · 256 GB/sceiling 15.38–11 stock24–36 tuned

Swipe sideways to see the whole chart →

Sources: ceiling computed by this site's engine. Stock: Hardware Corner (llama.cpp, MTP off, 4k). Tuned: syv-ai (RTX 3090) and one aggregator (other rows) — community-reported, not measured by us. Checked 2026-10-08.
In words: stock llama.cpp runs Qwen3.8-27B at 40.3 (RTX 3090), 46.2 (RTX 4090), 74.8 (RTX 5090) and 12.2 (DGX Spark) tokens per second; tuned stacks reach 127, 148 and 33–35.

For a dense model each token reads every weight once, so speed is capped at memory bandwidth divided by the weights: about 55.8 tokens per second on an RTX 3090 and 106.8 on an RTX 5090, by our engine. Stock llama.cpp lands at 70 to 77 per cent of the ceiling on the machines Hardware Corner measured. Speculative decoding bends the rule: the syv-ai stack reached 127 tokens per second on one RTX 3090 and about 1,035 across 64 requests (repo). The RTX 5090, DGX Spark and Strix Halo (24–36 tuned, 8–11 stock) tuned figures come from one aggregator, Context Studios. We did not run any of these.

That is also why the 128 GB tier feels slower than its size suggests: unified-memory boxes have a fraction of a graphics card's bandwidth. Mixture-of-experts picks — K2 Horizon MoVA, Qwen3.8-Flash-Next — read only their active parameters per token, which is how they stay usable.

Going fully offline

The fully offline coding stack

How a fully offline local coding setup is wiredFive connected boxes: a coding agent such as Claude Code, a local API on port 11434, a runtime such as Ollama or llama.cpp, model weights downloaded once (the only step that needs the internet), and your GPU or unified memory. Below, three settings: a context of at least 64,000 tokens, reasoning effort at medium, and planning for web tools not working offline.The fully offline coding stackFive parts, one network dependency, three settings that make or break it.internet needed here, onceeverything else runs with the cable pulled1 AgentClaude Code, QwenCode, OpenCode,Cline, Aider2 Local APIlocalhost:11434,Anthropic- orOpenAI-compatible3 RuntimeOllama, llama.cpp,LM Studio, vLLM,MLX4 WeightsGGUF or MLX file,downloaded once5 HardwareGPU VRAM or unifiedmemory: weights +KV cacheTHREE SETTINGS THAT DECIDE WHETHER IT WORKS1. Context ≥ 64,000Ollama defaults to 4k below 24 GiB ofVRAM.OLLAMA_CONTEXT_LENGTH=64000 ollama serve2. Effort: start at mediumQwen3.8-27B defaults to xhigh andoverthinks. Time whole tasks.reasoning_effort: medium3. Expect no web toolsWeb search and fetch need the network.Plan for it.CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1

Swipe sideways to see the whole chart →

Source: Ollama documentation (context length; Claude Code integration). Checked 2026-10-08. Run it once with the network disconnected before you rely on it.
In words: five parts — agent, local API, runtime, weights, hardware — and only the one-time download needs the internet.

After the one-time download, nothing needs the network. That rules out web search and fetch tools, package installs without a local mirror, and anything that signs in to a service.

# 1. One-time, online: install Ollama and pull a model from your tier.
ollama pull qwen3.8:27b

# 2. Serve with a context long enough for a coding agent.
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

# 3. Point Claude Code at the local server, and switch off its non-essential traffic.
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
claude --model qwen3.8:27b        # or: ollama launch claude

The variables are the ones in Ollama's Claude Code guide, checked 2026-10-08. Run ollama ps and check that the CONTEXT column shows what you asked for: Ollama defaults to 4k below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k above. Ollama's Anthropic-compatible API does not support prompt caching, so long sessions re-read their context; expect slower prompt processing than a hosted run. Qwen Code, OpenCode, Cline, Aider and Continue all work against the same server — see the coding-agents guide. Then unplug the network cable and run a real task: anything that quietly needed the internet shows up in ten minutes. More in offline AI.

Honest limits and methodology

  • We did not run these models. Speeds are Hardware Corner's (stock), the syv-ai repository's (tuned RTX 3090) and one aggregator's (other tuned rows). Memory totals, fit verdicts and bandwidth ceilings are this site's engine; Hardware Corner's measurements agree with it.
  • The independent index is general, not coding-only. Dedicated coders like Devstral Small 2 and Qwen3-Coder-Next score low on it or are not scored, and can still be good at code. That is why vendor coding scores sit beside it, labelled.
  • K2 Horizon is new and lightly tested. Strong independent scores, but no Ollama tag yet and fresh llama.cpp support. If you need it working today with no build steps, take each tier's Ollama pick.
  • Quantisation costs quality. Every score here is for the full-precision model; nobody has benchmarked the 4-bit builds you will run, and the 3-bit DeepSeek is a particular worry.
  • Benchmarks are not your repository. Run your last three real tasks through a local setup and count how many you would have accepted.
Sources. Independent: Artificial Analysis Intelligence Index v4.3.2 and Coding Agent Index v1.5, METR, Simon Willison. Vendor-reported: model cards from Alibaba (Qwen), MBZUAI IFM (K2 Horizon), Mistral, Anthropic. Sizes: Ollama library and official GGUF listings. Attention layouts for Qwen3.8-27B, Qwen3.5-4B/9B, Qwen3.6-35B-A3B, Qwen3.8-Flash-Next and Devstral Small 2 were read from their published configs for this report. All checked 2026-10-08; the formulas are on the methodology page. AA marks some rows with an asterisk it does not explain (Gemma 4 12B, Qwen3.5 4B among them); we leave those out. Not covered: hardware prices (we link to current listings instead), autocomplete models, fine-tuning, multi-GPU rigs, and anything released after 2026-10-08. Related: Report #2, the July snapshot of open coding models · best local coding models, the evergreen guide · best local LLM for coding · VRAM for coding models · local vs cloud for coding

Frequently Asked Questions

What is the best local coding model right now?

It depends on your memory. On 8–16 GB, K2 Horizon 7B (Artificial Analysis index 21). On a 24–32 GB GPU, Qwen3.8-27B (34). On a 128 GB machine, Qwen3.8-Flash-Next (40), though it is a tight fit. Qwen3.8-27B is the best overall balance of quality and ease of running.

What is the best coding model for an 8 GB GPU or laptop?

K2 Horizon 3.7B, a 3.2 GB file that scores 16 on the independent index. If you want an Ollama one-liner instead, Qwen3.5-4B is a 3.3–4.0 GB download. Use either for small, well-specified edits rather than multi-file agent work.

What is the best coding model for a 12 or 16 GB GPU?

K2 Horizon 7B: a 5.6 GB file scoring 21 on the independent index, nearly double Qwen3.5-9B's 11. It needs a recent llama.cpp build because Ollama has no tag for it yet. The easiest alternative is Qwen3.5-9B, installed with ollama pull qwen3.5:9b.

What is the best coding model for an RTX 3090, 4090 or 5090?

Qwen3.8-27B. It scores 34 on the independent index, the highest of any model that fits one consumer GPU, and needs about 21.9 GB at 64k context. K2 Horizon MoVA 36B-A4B is a faster runner-up on a 32 GB RTX 5090, and Devstral Small 2 leaves the most room for long context on a 24 GB card.

What is the best coding model for a 128 GB DGX Spark, Strix Halo or Mac?

Qwen3.8-Flash-Next scores 40, the best of anything in reach, but its 120 GB build is tight on a DGX Spark or Strix Halo and too big for a 128 GB Mac at default settings. Qwen3.8-27B at 8-bit with full context is the comfortable choice on all three.

How much VRAM do I need for Qwen3.8-27B?

About 17.8 GB at short context, 21.9 GB at 64k and 26.2 GB at 128k, at Q4_K_M with a 16-bit KV cache, by this site's engine; Hardware Corner measured 18, 22 and 26 GB. A 24 GB GPU handles 64k and a 32 GB GPU handles 128k.

How good are local coding models compared with Claude, GPT and Gemini?

On the independent Artificial Analysis index, the best model that fits one GPU scores 34 and the top subscription models score 53 to 58. That matches the budget tier, such as Claude Haiku 5.5 at medium effort, not the flagships.

Can I use Claude Code with a local model, fully offline?

Yes. Pull the model once with Ollama, then set ANTHROPIC_BASE_URL to http://localhost:11434, ANTHROPIC_AUTH_TOKEN to ollama and ANTHROPIC_API_KEY to an empty string. Set the context to at least 64,000 tokens. Web search and fetch tools will not work offline.

Why does my local coding agent forget the task halfway through?

Usually context. Ollama defaults to 4k tokens below 24 GiB of VRAM, while coding tools need at least 64,000. Set OLLAMA_CONTEXT_LENGTH=64000, restart the server and confirm with ollama ps.

Will a local model replace my coding subscription?

For some work, yes; for hard problems, not yet. The independent gap between the best home model and the best subscription model is 24 points. Most people do best with a hybrid: local for private and bulk work, a subscription for the problems that stall.