Best Local LLMs for Coding
Written by Jakub Rusinowski · Last updated September 11, 2026
Writing, refactoring and debugging code in an editor or terminal, with the model reading real project files.
Want one agentic coder and one autocomplete model matched to your exact GPU, with setup steps? Pick by memory tier →
Top pick: Qwen3-Coder 8B
Scores 96.3/100 for coding and software engineering. 8B parameters, needing about 5.6 GB at Q4_K_M, 125K context, Apache-2.0.
Ranked for coding and software engineering
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Qwen3-Coder 8B | 96.3 | 8B | 125K | Apache-2.0 | — (estimated) |
| 2. Qwen 3.7 35B-A3B | 95.9 | 35B | 256K | Apache-2.0 | — (estimated) |
| 3. Qwen 3 32B | 95.4 | 33B | 125K | Apache 2.0 | — (estimated) |
| 4. Qwen 3.6 35B-A3B | 95.2 | 35B | 256K | Apache-2.0 | — (estimated) |
| 5. Devstral Small 2 24B | 94.9 | 24B | 256K | Apache-2.0 | — (estimated) |
| 6. DeepSeek R1 Distill Qwen 32B | 94.9 | 32B | 128K | MIT | 87 (cited) |
Best pick for your memory budget
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Qwen3-Coder 8B (96.3) GLM-6 9B (93.2) Qwen 3 8B (92.9) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (95.9) Qwen3-Coder 8B (94.9) DeepSeek R1 Distill Qwen 14B (94.2) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Devstral Small 2 24B (96.5) Qwen 3 14B (95.9) DeepSeek R1 Distill Qwen 14B (93.9) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Qwen 3.7 35B-A3B (98.5) Qwen 3 32B (98) Qwen 3.6 35B-A3B (97.8) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Qwen 3.7 35B-A3B (96.9) Qwen 3.5 72B (96.6) Qwen 3.6 35B-A3B (96.2) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Devstral-2 123B (99) Qwen3-Coder-Next (80B-A3B MoE) (96.6) Qwen 3.5 122B-A10B (MoE) (96.5) |
How this ranking works
Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.
Worked example — Qwen3-Coder 8B: capability 88.4 × 0.419, quality 82.3 × 0.247, context 97.3 × 0.16, license 100 × 0.044, accessibility 100 × 0.129 + 6 tag bonus (coding, debugging, code-review).
Requirements applied: context floor 16,384 tokens (ideal 131,072), quality floor 55, licence weight 0.4, latency weight 0.7.
Running coding and software engineering locally
- Prompt processing, not decode, dominates perceived latency when a coding agent re-sends a large file on every turn.
- A smaller model at a higher quantization usually beats a larger model forced to Q3 for code, because quantization damage shows up first as syntax and API errors.
FAQ
What is the best local LLM for coding and software engineering?
Qwen3-Coder 8B, scoring 96.3/100 against this workload's published requirements. 159 models qualified.
What hardware do I need for coding and software engineering?
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
How were these models ranked?
Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.
Hardware for This Workload
- Best GPU for coding
- Best models for the NVIDIA GeForce RTX 4060 Ti 8GB
- Best models for the NVIDIA GeForce RTX 3080 Ti
- Best models for the NVIDIA GeForce RTX 4090 Laptop GPU
Related Workloads
- Best local LLMs for ocr
- Best local LLMs for agents
- Best local LLMs for cybersecurity
- Best local LLMs for document analysis