Local AI Coding Assistant: Run Your Own Copilot
A local AI coding assistant is three choices: a model (Qwen3-Coder, Devstral-2, Qwen 3.6 — picked by your VRAM), a tool (Continue for Copilot-style assist, Cline or Aider for autonomous agents, Tabby for teams), and the hardware to serve it. This hub covers all three, with every hardware figure computed by the same engine as our GPU checker.
Written by Jakub Rusinowski · Last updated August 4, 2026 · Hardware figures computed by our VRAM engine
Coding is the local-AI use case where the argument is strongest — code is the data companies least want leaving the building, coding assistants are subscription products you can actually replace, and 2026's open coding models are good enough that the replacement isn't a downgrade for everyday work.
It's also where hardware questions get sharpest. An autocomplete model must answer in milliseconds; an agentic model must hold a repo's worth of context; both must fit your VRAM at once. The guides below are organized so you can enter anywhere — by model, by tool, or by hardware — and each one links the others.
One cluster deserves calling out. Beyond picking a model and a tool there is the layer everyone discovers second: the loop the agent runs in, and the harness that runs it. On a frontier API you can get away with ignoring both. On your own GPU they are where most of the quality is — a verifier command, a raised context window and a rules file routinely move results further than a model-tier upgrade, and they cost no VRAM at all. That's the Agent engineering section.
The stack in 30 seconds
One model per hardware tier — full rankings in the model guide below. VRAM figures come from the same engine as the compatibility checker.
| Model | VRAM (Q4) | Runs on | Context | License |
|---|---|---|---|---|
| Qwen3-Coder 8B Entry (8 GB GPU) Current-gen coder + autocomplete on any modern GPU or 16 GB Mac. ollama pull qwen3-coder:8b | 5.6 GB | 8 GB GPU (RTX 3060/4060) Mac: 16 GB unified | 125K | Apache-2.0 |
| Devstral 2 22B Mid (16 GB GPU) The best dedicated coding agent under 24 GB. Apache 2.0. ollama pull devstral:22b | 14.1 GB | 16 GB GPU (RTX 4060 Ti 16GB / 5060 Ti) Mac: 24 GB unified | 125K | Apache-2.0 |
| Qwen 3.6 27B Sweet spot (24 GB GPU) The realistic daily driver for local agentic coding on a 3090/4090. ollama pull qwen3.6:27b | 17.6 GB | 24 GB GPU (RTX 3090/4090) Mac: 24 GB unified | 256K | Apache-2.0 |
| Qwen3-Coder-Next (80B-A3B MoE) Workstation / Mac 96 GB+ ~96% of the open flagship’s quality, self-hostable, fast MoE. ollama pull qwen3-coder:80b-a3b-q4 | 49.1 GB | 2×48 GB GPUs / big unified memory Mac: 96 GB unified | 256K | Apache-2.0 |
The guides (21)
- Best Coding LLM for 8GB, 16GB & 24GB VRAM (2026)Every serious open coding model, ranked by the VRAM tier that actually runs it.
- Run a Local Coding Agent End-to-End: Ollama + Qwen3-Coder + VS CodeOllama + Qwen3-Coder + Continue/Cline in VS Code, end to end, in ~40 minutes.
- Local GitHub Copilot Alternatives: What Actually Replaces It OfflineContinue, Tabby, Cline, Aider, Twinny — what each replaces, fully offline vs hybrid.
- Continue vs Cline vs Aider vs Tabby: the Local Coding Tool MatrixThe feature matrix: assistant vs agent vs terminal vs team server — and sensible pairings.
- VRAM Requirements for Coding LLMs: What Every Tier Actually RunsWhat 8, 12, 16, 24 and 48+ GB actually run — with the context-headroom rule that generic tables miss.
- Best GPU for Local AI Coding in 2026: Five Picks That Make SenseFive GPU picks matched to the coding models each one runs — VRAM first, compute second.
- Local Coding Agent vs Copilot vs Cursor: What Each Really CostsLocal vs Copilot ($10/mo) vs Cursor ($20/mo) vs metered APIs — honest math, including where local loses.
- Why Your Local Coding Agent Feels Slow — and the Six Fixes That WorkAgent turns are prefill-bound: prompt-cache reuse, speculative decoding and the fixes that beat a GPU upgrade.
- Loop Engineering on a Local LLM: Agent Loops That Check Their Own WorkThe act–observe–verify cycle: four loop levels, verifier design, iteration caps and the failure modes.
- The Agent Harness: What Actually Turns a Local Model Into a Coding AgentWhat actually turns a model into an agent — the six jobs a harness does, and how to tune each locally.
- Context Engineering for Local Coding Agents: Budgeting a Window You Pay ForBudgeting a window you pay for in VRAM: what fills it, when to compact, and beating context rot.
- Tool Calling With Local Models: Making Agents Stop Breaking on Malformed JSONWhy local agents break on malformed JSON — chat templates, GBNF grammars and repair strategies.
- AGENTS.md for Local Coding Agents: The Cheapest Quality Upgrade You Can MakeThe rules file that fixes small-model agents: what to write, what to cut, AGENTS.md vs CLAUDE.md.
- MCP With Local Coding Agents: Which Servers Earn Their TokensWhich MCP servers earn their tokens on a 32K window — and when a shell command beats a server.
- Sandboxing a Local Coding Agent: Autonomy Without the Blast RadiusDevcontainers, throwaway worktrees, permission allowlists and egress rules for unattended runs.
- Evaluating a Local Coding Agent: Benchmarks Lie, Your Repo Doesn’tBenchmarks shortlist; your repo decides. Build a 20-task eval from your own git history.
- A Local Code Review Agent: The Use Case Small Models Are Actually Good AtThe best quality-per-VRAM use case: read-only review, in a pre-commit hook or self-hosted CI.
- Run Qwen3-Coder Locally: 8B, 80B-A3B, and the 480B QuestionThe 8B, the 80B-A3B, the expert-offload trick, and when to just use the API for the 480B.
- Aider with Local Models: Setup and the Models That Actually WorkThe terminal pair programmer on Ollama — config, context, and the models that follow its edit formats.
- Cline with Local Models: What Actually Works (and What Doesn’t)The VS Code agent on local models — the context-window fix and the honest capability tiers.
- Terminal Coding Agents on Local Models: OpenCode, Qwen Code, Aider & FriendsOpenCode, Qwen Code and Aider pointed at your own GPU — config, model picks and honest limits.
Coding models move fast — the biweekly digest covers each new release with the same hardware-first treatment, and every guide above carries a last-updated date.
Frequently asked questions
Can I really replace GitHub Copilot with a local model?
For autocomplete, chat, and inline edits — yes: Continue.dev plus a FIM-tuned local model (Qwen3-Coder, StarCoder2) reproduces the Copilot experience offline, and on a 16–24 GB GPU the quality gap has become marginal for mainstream languages. Frontier cloud models still lead on the hardest multi-file agent tasks, which is why many developers run local-first with an API fallback.
What hardware do I need for a local coding assistant?
Entry: any 8 GB GPU or a 16 GB Apple Silicon Mac runs StarCoder 2 7B for completion and small edits. Sweet spot: a 24 GB card (RTX 3090/4090) runs Qwen 3.6 27B — the tier where local agentic coding gets genuinely good. Workstation: 2×24 GB or a 96–128 GB Mac runs Qwen3-Coder 80B-A3B. Check your exact GPU with our free compatibility checker.
Which tool should I use: Continue, Cline, Aider, or Tabby?
Continue if you want a Copilot replacement (autocomplete + chat in VS Code/JetBrains). Cline if you want an autonomous agent that plans and edits multiple files. Aider if you live in the terminal and want git-native agent workflows. Tabby if you are provisioning one self-hosted server for a whole team. They are not mutually exclusive — Continue + Cline is a common pairing.
What is the best local coding model right now?
Per hardware tier: StarCoder 2 7B on 8 GB, Devstral Small 2 24B on 16 GB, Qwen 3.6 27B or Qwen 2.5 Coder 32B on 24 GB, and Qwen3-Coder 80B-A3B at workstation scale. The open frontier (DeepSeek V4-Pro, GLM, Kimi K2) is stronger still but needs datacenter hardware — see our dated Report #2 for the full landscape.
What is loop engineering, and why does it matter more for local models?
Loop engineering is designing the cycle an agent runs in — act, observe the result, check it against an objective condition such as a test suite or type check, then retry or stop — instead of prompting step by step. It matters disproportionately on local hardware because the deterministic check supplies the judgement a 22–32B model lacks, and local iterations cost time rather than money. Adding one test-command gate to an existing setup is usually a larger quality jump than moving up a model tier.
What is an agent harness?
The software layer around the model that makes it an agent: it assembles the prompt, parses output into tool calls, executes them, compacts context as the window fills, enforces permissions and persists state. The same open weights behave very differently under different harnesses, which is why model cards state which one produced their benchmark score. Locally the harness is the half you can still improve after your VRAM is fixed.
Is a local coding assistant cheaper than Copilot or Cursor?
If you already own a capable GPU — immediately: electricity for all-day coding inference is a few dollars a month versus $10/month (Copilot Pro) or $20/month (Cursor Pro) indefinitely. If you are buying hardware for it, a used RTX 3090 pays back against a Cursor subscription in roughly three years — so buy for the privacy and control, and treat the subscription savings as a bonus.