// Coding agents

Best Coding LLM for 8GB, 16GB & 24GB VRAM (2026)

On a 24 GB card (RTX 3090/4090), run Qwen 3.6 27B (~17.6 GB at Q4_K_M), or Devstral Small 2 24B (~15.3 GB) if you want an agentic specialist. With a 64 GB-class machine, Qwen3-Coder-Next (80B total, 3B active) is the most capable coder here that one machine can hold: all 80B parameters must fit, but only 3B are read per token, so it decodes fast. On 8 GB, use StarCoder 2 7B for autocomplete (~5.1 GB) — agentic work starts at 16 GB and is comfortable at 24 GB. Every VRAM figure below comes from our compute module, not marketing pages.

Written by Jakub Rusinowski · Last updated September 29, 2026 · Hardware figures computed by our VRAM engine

Most "best coding LLM" lists rank models nobody can run. This one is sorted the other way around: by the hardware tier you actually have. Each table shows the model's real Q4_K_M VRAM requirement — the same math behind our GPU compatibility checker — plus the GPU tier and the smallest Mac unified-memory config that fits it.

Two things to know before the tables. First, coding models split into agentic coders (built to drive an editing loop in Continue, Cline, or Aider) and autocomplete models (small, fill-in-the-middle-tuned, judged on latency). Serious local setups run one of each. Second, the frontier moves fast, and published coding benchmarks disagree between sources and scaffolds, so this page ranks by what fits your hardware rather than quoting scores. For a dated, citable snapshot of the whole landscape, see our Local AI Report #2: the best open-source coding models. If you want every local model scored and ranked for coding rather than sorted by tier, see the full benchmark ranking →.

The best coders you can download (workstation tier)

Open-weights state of the art. These need multi-GPU rigs or big unified memory all-in-VRAM — but note the MoE offload trick below the table.

ModelVRAM (Q4)Runs onContextLicense
Qwen3-Coder 480B-A35B (MoE)
Open state of the art
The largest dedicated open coder from Qwen, which positions it near Claude Sonnet 4 on agentic coding. Realistically an API or datacenter model.
ollama pull qwen3-coder:480b-a35b
290.6 GBMulti-GPU server — or use an API
Mac: 512 GB unified
256KApache-2.0
Devstral-2 123B
Agentic specialist
Mistral’s agentic flagship, built for multi-file edits and repo navigation. A dense 123B, so every parameter is read on every token: slower than the MoE options here.
ollama pull devstral:123b
75.1 GB2×48 GB GPUs / big unified memory
Mac: 128 GB unified
256KApache-2.0
Qwen3-Coder-Next (80B-A3B MoE)
Best self-hostable coder
Released as Qwen3-Coder-Next: 80B total but only 3B active per token, on a hybrid-attention base whose long-context cache is cheap. All 80B must be resident, so plan for a 64 GB-class machine.
ollama pull qwen3-coder:80b-a3b-q4
49.1 GB2×48 GB GPUs / big unified memory
Mac: 96 GB unified
256KApache-2.0
The MoE offload trick. Small-active-parameter MoE models only need their active experts on the GPU. With llama.cpp's expert-offload options, the inactive experts can sit in system RAM while the active ones stay in VRAM. It is slower than the all-in-VRAM figure in the table, but it turns a "workstation model" into something a well-RAMed gaming PC can use. The table shows the all-in-VRAM (fastest) case.

Best coding models for 16–24 GB GPUs

The consumer sweet spot — RTX 3090/4090, RX 7900 XTX, or a 32 GB Mac. This is where "self-hosted and genuinely good at code" starts.

ModelVRAM (Q4)Runs onContextLicense
Qwen 3.6 27B
Best on a 24 GB card
The strongest all-round coder that fits an RTX 3090/4090 comfortably at Q4 — the realistic daily driver for local agentic work.
ollama pull qwen3.6:27b
17.6 GB24 GB GPU (RTX 3090/4090)
Mac: 24 GB unified
256KApache-2.0
Qwen3-Coder 30B-A3B (MoE)
Fast MoE coder
A dedicated coder with 3.3B active parameters, so it streams like a small model while holding 30B of knowledge. Fits a 24 GB card at Q4.
ollama pull qwen3-coder:30b
19.2 GB24 GB GPU (RTX 3090/4090)
Mac: 32 GB unified
256KApache-2.0
Devstral Small 2 24B
Agentic specialist
Mistral’s small agentic coder, built to drive an editing loop. A dense 24B that fits a 24 GB card with room for context; on 16 GB it is a tight fit.
ollama pull devstral-small-2:24b
15.3 GB24 GB GPU (RTX 3090/4090)
Mac: 24 GB unified
256KApache-2.0
Codestral 22B
FIM autocomplete (license warning)
Excellent fill-in-the-middle quality across 80+ languages — but the MNPL license forbids commercial/production use. Hobby projects only.
ollama pull codestral:22b
14.2 GB16 GB GPU (RTX 4060 Ti 16GB / 5060 Ti)
Mac: 24 GB unified
32KMNPL-0.1 (non-production)
License check before you standardize on a model. The License column in each table comes from the model's own record — read it before you build on a model. Codestral 22B is MNPL (non-production): fine to evaluate at home, not fine inside a company workflow. StarCoder2 ships under BigCode OpenRAIL-M, which permits commercial use with use-based restrictions. If you’re picking a model for work, this column matters as much as any benchmark score.

Small coding models (4–12 GB) and autocomplete

Latency beats size for inline completions — a FIM-tuned small model next to your cursor outperforms a big chat model that takes two seconds to answer.

ModelVRAM (Q4)Runs onContextLicense
StarCoder 2 7B
Best small autocomplete
A fill-in-the-middle code model that fits any 8 GB GPU — the default pick for autocomplete on modest hardware.
ollama pull starcoder2:7b
5.1 GB8 GB GPU (RTX 3060/4060)
Mac: 16 GB unified
16KBigCode OpenRAIL-M
StarCoder 2 15B
Mid-size autocomplete
A 600-language code model for 12 GB cards, for when you want more breadth than the 7B.
ollama pull starcoder2:15b
10.2 GB12 GB GPU (RTX 3060 12GB / 4070)
Mac: 16 GB unified
16KBigCode OpenRAIL-M
StarCoder 2 3B
Fastest autocomplete
The lightweight option: small enough that ghost text keeps up with your typing. Pair it with a bigger chat model.
ollama pull starcoder2:3b
2.6 GB8 GB GPU (RTX 3060/4060)
Mac: 16 GB unified
16KBigCode OpenRAIL-M

The frontier you can’t download (and when to call it)

The largest open-weight coders — DeepSeek V4, GLM-5.x, and Kimi K2.x — want 80 GB-class multi-GPU hardware. For a solo developer they’re API models, not local models. The pattern that wins: a fast local coder for the inner loop (autocomplete, quick edits, private code) plus a frontier model over an API for the hardest problems. Our cloud AI directory lists where each frontier model is hosted and what it costs per token.

How to choose from these tables. Start from your VRAM, not from the leaderboard: pick the largest agentic coder your tier runs comfortably, add StarCoder 2 3B or StarCoder 2 7B for autocomplete if your card has headroom, and check the exact fit for your GPU — including quant options and context-length headroom — with the compatibility checker. Then wire it into an editor with our end-to-end setup guide.

The hardware these tiers assume

The 24 GB tier is where local coding gets genuinely good, and two cards own it: the RTX 4090, and the used RTX 3090 as the budget route to the same VRAM. A 16 GB card such as the RTX 5060 Ti 16GB runs a mid-size coder or an autocomplete model plus a small chat model, but not Devstral Small 2 and an autocomplete model side by side. Street prices are well above launch MSRP right now, so compare before you buy.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
NVIDIA GeForce RTX 4090 24GB
24 GB VRAM · 450 W board power
2026 prices are volatile — check the current listing.
NVIDIA GeForce RTX 3090 24GB
24 GB VRAM · 350 W board power
2026 prices are volatile — check the current listing.
NVIDIA GeForce RTX 5060 Ti 16GB
16 GB VRAM · 180 W board power
2026 prices are volatile — check the current listing.

No hardware? Rent the GPU first

Hardware falls short of the model you want? Rent a 24–48 GB GPU by the hour for a few dollars and test-drive the exact model before buying anything.

  • RunPod — RTX 4090 ≈ $0.3–0.7/hr, A100/H100 by the hour; serverless per-second billing available
  • Vast.ai — Marketplace of hosted GPUs — often the cheapest per hour (interruptible options)
  • Lambda — On-demand A100/H100/B200 instances and clusters, per-hour billing

Affiliate links — we may earn a commission if you sign up, at no extra cost to you.

Full list on the cloud AI directory.

Frequently asked questions

What is the best local coding LLM for a 24 GB GPU (RTX 3090/4090)?

Qwen 3.6 27B is the best all-round coder that fits a 24 GB card at Q4 (~17.6 GB). Devstral Small 2 24B (~15.3 GB) is the agentic specialist, and Qwen3-Coder 30B-A3B (~19.2 GB) is the fast MoE option. Run one of them for agentic work and StarCoder 2 for autocomplete if you have the headroom.

What is the best coding model for 16 GB VRAM?

Devstral Small 2 24B needs ~15.3 GB at Q4 before the KV cache, so on a 16 GB card it fits only with a short context. For comfortable headroom, run StarCoder 2 15B (~10.2 GB) or StarCoder 2 7B (~5.1 GB) for autocomplete, and use the compatibility checker to test any agentic model against your exact card and context length.

Can I run Qwen3-Coder locally?

Yes, at three sizes. The 30B-A3B MoE (~19.2 GB at Q4) fits a 24 GB card. Qwen3-Coder-Next (80B-A3B) needs ~49.1 GB all-in-VRAM — a 64 GB-class machine — because every expert must be resident even though only 3B are active per token. The 480B-A35B flagship needs ~290.6 GB — datacenter or API territory.

Are local coding models as good as GitHub Copilot or Cursor?

For autocomplete and everyday edits on a 24 GB card, local models are close enough for many developers. For the hardest multi-file agentic tasks, frontier API models still lead. The winning setup is usually local-first with an API escape hatch.

Which local coding models allow commercial use?

Check the License column in each table, which comes from the model’s own record. Qwen3-Coder and Qwen 3.6 ship under Apache 2.0. Codestral 22B uses Mistral’s non-production license, so it is not for commercial work, and StarCoder2 uses BigCode OpenRAIL-M (commercial use with use-based restrictions).