Local AI Report #2 — The Best Open-Source Coding Models Right Now
A field guide to the strongest open-weight coding models in mid-2026: the SWE-bench frontier (DeepSeek V4-Pro, GLM-5.2, Kimi K2.6), the best coder you can actually download (Qwen3-Coder), and the setup that gives the most code per dollar on a single GPU.
- Frontier open coders now clear 70–80% on SWE-bench Verified — within a few points of the best closed models. The three families to know: DeepSeek V4-Pro, GLM-5.2 and Kimi K2.6.
- The best coder you can actually download is Qwen3-Coder: the 480B-A35B flagship is open state-of-the-art, and the 80B-A3B variant trims it to a single-workstation footprint at ~96% of the quality.
- Nothing at the frontier fits a consumer GPU — the top models want 80 GB-class multi-GPU or big unified memory. On a 24 GB card, Qwen 3.6 27B and Devstral-2 are the realistic self-hosted coders.
- Licenses finally line up with commercial use: GLM-5.x, Kimi K2.x and DeepSeek V4 ship under MIT; Qwen3-Coder and Devstral-2 under Apache 2.0.
- Benchmarks pick the shortlist; your repo picks the winner. The pattern that wins in 2026 is a fast local coder for the inner loop plus a frontier model over an API for the hard 5%.
New & Notable Models
| Model | Params | VRAM (Q4) | Notes |
|---|---|---|---|
| DeepSeek V4.1-Pro | 1.6T / 49B active | ~400 GB | Open SWE-bench Verified leader (~80%). Reasoning-grade coding — a datacenter node or, realistically, an API call. MIT-licensed. |
| GLM-5.2 | 744B / 40B active | ~400 GB | The strongest long-horizon agentic coder; tops the open field on SWE-bench Pro. MIT, with 1M-context builds. |
| Kimi K2.6 | 1T / 32B active | ~320 GB | Best-in-class MCP tool-use and agent-swarm runs; ~97% HumanEval. Built for long autonomous coding sessions. |
| Qwen3-Coder 480B-A35B | 480B / 35B active | ~270 GB | The best dedicated open coder — state-of-the-art open on SWE-bench Verified, which Qwen rates near Claude Sonnet 4. Apache 2.0. |
| Qwen3-Coder 80B-A3B | 80B / 3B active | ~48 GB | The frontier coder you can self-host: ~96% of flagship quality, and 3B active params keep tokens fast. Offload experts to cut VRAM. |
| Devstral-2 123B | 123B / 40B active | ~68 GB | Purpose-built agentic SWE model — 71.6% SWE-bench Verified. Apache 2.0 and easy to fine-tune on your own stack. |
| MiniMax M3 | 230B / 10B active | ~130 GB | Cost-efficient coder with a genuine 1M-token window; 59% on SWE-bench Pro. The pick for whole-repo context. |
| Qwen 3.6 27B | 27.8B | ~17 GB | The best coder that fits a 24 GB card at Q4 — the realistic daily driver for local agentic work. |
| StarCoder 2 7B | 8B | ~6 GB | Runs on almost anything. The fast inner-loop model for autocomplete, fill-in-the-middle and quick edits. |
Hardware Watch
Coding models split cleanly into "download and run" and "admire from an API." The dividing line is memory, and for the 2026 frontier it sits far above any single consumer card.
- Frontier tier (270–400 GB): DeepSeek V4.1-Pro, GLM-5.2 and Kimi K2.6 need 80 GB-class multi-GPU rigs or big unified memory. An Apple M3 Ultra (512 GB) is the one desk-side box that loads the biggest MoEs at Q4; otherwise these are API models.
- Workstation tier (48–130 GB): Qwen3-Coder 80B-A3B, Devstral-2 123B and MiniMax M3 fit a single Apple M4 Max (128 GB) or a 2×48 GB GPU pair. This is where "self-hosted and genuinely good at code" starts.
- The MoE-offload trick: small-active-parameter models — Qwen3-Coder 80B-A3B (3B active) and Qwen 3.6 35B-A3B (3B active) — only need the active experts resident in VRAM. With llama.cpp expert-offload the 80B-A3B runs on as little as ~8 GB VRAM plus 32 GB system RAM; the ~48 GB figure in the table is the all-in-VRAM (fastest) case.
- Single-GPU reality (16–24 GB): Qwen 3.6 27B at Q4 is the best coder that fits a NVIDIA GeForce RTX 3090 or NVIDIA GeForce RTX 4090 comfortably; Devstral Small 2 24B and StarCoder 2 7B cover a NVIDIA GeForce RTX 4060 Ti 16GB.
Tooling Updates
The model is half the story; the harness around it is the other half.
- Runtimes — Ollama and LM Studio ship tags for the consumer-sized coders (Qwen3-Coder 30B-A3B / 80B-A3B, Qwen 3.6, Devstral Small 2 24B); the 200 GB+ frontier models are llama.cpp / vLLM territory or cloud endpoints. Confirm your runtime version before pulling a new MoE — expert-offload support moves fast.
- Agentic harnesses — Aider, Continue, Cline and OpenCode all point at a local OpenAI-compatible endpoint (Ollama, LM Studio or vLLM). A model's SWE-bench number assumes an agent loop like these; a raw chat box leaves a lot of the score on the table.
- Tool use / MCP — Kimi K2.6 leads on MCP tool-calling, which matters more than raw generation for multi-file, multi-step edits. If your workflow is tool-heavy, weight tool-use benchmarks over HumanEval.
- Fill-in-the-middle — for inline autocomplete a FIM-tuned small model (StarCoder 2 7B, Codestral) beats a bigger chat model, and you can run it alongside your agentic model.
The Sweet Spot
- Inner loop (local): Qwen 3.6 27B at Q4 on a 24 GB card for chat, refactors and quick edits — fast, private and free per token. Add StarCoder 2 7B for fill-in-the-middle autocomplete.
- Agentic loop (local, if you have the memory): Qwen3-Coder 80B-A3B with expert-offload, driven by Aider or Cline, handles most repo-level tasks at a fraction of frontier VRAM.
- The hard 5% (API): route the rare task that stumps the local model to GLM-5.2 or DeepSeek V4.1-Pro over an API. Paying per token for the tasks that need it beats buying a 400 GB rig you use at 5% duty.
This two-model pattern — a cheap local coder for the inner loop, a frontier model on tap for the hard cases — is the setup that keeps winning. The cost calculator makes the local-vs-API break-even explicit for your token volume.
Workshop Note
From recent workshop sessions, three things practitioners keep re-learning. SWE-bench Verified is Python-heavy — a model that tops it can still trail on your Rust or TypeScript repo, so run a small eval on your own codebase before committing. The scaffolding is worth as many points as the model — the same weights score very differently under a good agent loop versus a bare prompt, so invest in the harness (Aider/Cline config, test hooks, MCP tools) before chasing a bigger model. Tighter context beats bigger context — feeding a whole repo just because a model advertises a 1M-token window usually hurts both accuracy and latency; retrieve the files that matter. Start from the smallest model that clears your own eval, then scale up only where it actually fails.
The October update: Report #6 — the October update: three best coding models for each hardware tier