Written by Jakub Rusinowski · Last updated June 26, 2026
Writing, refactoring and debugging code in an editor or terminal, with the model reading real project files.
Top pick: Kimi K2.5
Scores 97.2/100 for coding and software engineering. 32B parameters, needing about 20.1 GB at Q4_K_M, 125K context, Kimi License (research).
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Kimi K2.5 | 97.2 | 32B | 125K | Kimi License (research) | — (estimated) |
| 2. Qwen3-Coder 8B | 96.3 | 8B | 125K | Apache-2.0 | — (estimated) |
| 3. Qwen 3.7 35B-A3B | 95.9 | 35B | 256K | Apache-2.0 | — (estimated) |
| 4. Qwen 3 32B | 95.4 | 33B | 125K | Apache 2.0 | — (estimated) |
| 5. Qwen 3.6 35B-A3B | 95.2 | 35B | 256K | Apache-2.0 | — (estimated) |
| 6. DeepSeek R1 Distill Qwen 32B | 94.9 | 32B | 128K | MIT | 87 (cited) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Qwen3-Coder 8B (96.3) GLM-6 9B (93.2) Qwen 3 8B (92.9) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (95.9) Qwen3-Coder 8B (94.9) DeepSeek R1 Distill Qwen 14B (94.2) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3 14B (95.9) Devstral-2 22B (94) DeepSeek R1 Distill Qwen 14B (93.9) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Kimi K2.5 (99.7) Qwen 3.7 35B-A3B (98.5) Qwen 3 32B (98) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Kimi K2.5 (97.7) Qwen 2.5 72B Instruct (97.6) Qwen 3.7 35B-A3B (96.9) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Devstral-2 123B (98.5) Qwen 3.5 122B-A10B (MoE) (96.5) Qwen3-Coder 80B-A3B (MoE) (96.2) |
Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.
Worked example — Kimi K2.5: capability 94.4 × 0.419, quality 91.3 × 0.247, context 97.3 × 0.16, license 70 × 0.044, accessibility 80 × 0.129 + 6 tag bonus (coding, debugging).
Requirements applied: context floor 16,384 tokens (ideal 131,072), quality floor 55, licence weight 0.4, latency weight 0.7.
Kimi K2.5, scoring 97.2/100 against this workload's published requirements. 110 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.