Best Local LLMs for Agents
Written by Jakub Rusinowski · Last updated October 6, 2026
Multi-step autonomous loops where the model calls tools, reads results and decides what to do next.
Top pick: Ternary Bonsai 27B
Scores 91.7/100 for autonomous agents and tool use. 27B parameters, needing about 17.1 GB at Q4_K_M, 256K context, Apache 2.0.
Ranked for autonomous agents and tool use
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Ternary Bonsai 27B | 91.7 | 27B | 256K | Apache 2.0 | — (estimated) |
| 2. Devstral Small 2 24B | 91.2 | 24B | 256K | Apache-2.0 | — (estimated) |
| 3. Nemotron-Cascade 2 30B-A3B | 90.6 | 32B | 1M | NVIDIA Open Model License | — (estimated) |
| 4. Nemotron 3.5 Lightning 30B-A3B | 90.3 | 30B | 977K | OpenMDW-1.1 | — (estimated) |
| 5. Qwen3-Coder 30B-A3B (MoE) | 89.9 | 31B | 256K | Apache-2.0 | — (estimated) |
| 6. Seed-OSS-36B-Instruct | 89.5 | 36B | 512K | Apache-2.0 | — (estimated) |
Best pick for your memory budget
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-6 9B (86.8) GLM-5 9B (85.4) K2 Horizon 7B (78.6) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | GLM-6 9B (86) GLM-5 9B (84.6) Qwen 3.5 14B (84.6) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Devstral Small 2 24B (92.6) GLM-6 9B (84.7) Qwen 3.5 14B (84.4) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Ternary Bonsai 27B (94) Nemotron-Cascade 2 30B-A3B (92.8) Devstral Small 2 24B (92.6) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Ternary Bonsai 27B (91.6) Nemotron-Cascade 2 30B-A3B (91) Nemotron 3.5 Lightning 30B-A3B (90.6) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Devstral-2 123B (95.2) Qwen3.8-Flash-Next (93.9) Qwen3-Coder-Next (80B-A3B MoE) (93.1) |
How this ranking works
Reliable structured tool calls are a hard gate here, not a bonus: a model with no tool-use tag is excluded regardless of score. Coding weight is high (40%) because tool arguments are effectively code. Context needs match RAG — an agent loop accumulates every observation it has ever seen.
Worked example — Ternary Bonsai 27B: capability 89.3 × 0.427, quality 87.7 × 0.251, context 100 × 0.163, license 70 × 0.045, accessibility 80 × 0.113 + 3 tag bonus (agents).
Requirements applied: context floor 32,768 tokens (ideal 262,144), quality floor 60, licence weight 0.4, latency weight 0.6. Tool-calling capability is required.
Running autonomous agents and tool use locally
- Agent context grows monotonically across a run; a 32K window that looks generous at step 1 is exhausted by step 20.
- A single malformed tool call can end a run, so consistency matters more than peak capability.
FAQ
What is the best local LLM for autonomous agents and tool use?
Ternary Bonsai 27B, scoring 91.7/100 against this workload's published requirements. 55 models qualified.
What hardware do I need for autonomous agents and tool use?
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
How were these models ranked?
Reliable structured tool calls are a hard gate here, not a bonus: a model with no tool-use tag is excluded regardless of score. Coding weight is high (40%) because tool arguments are effectively code. Context needs match RAG — an agent loop accumulates every observation it has ever seen.
Hardware for This Workload
- Best GPU for agents
- Best models for the NVIDIA GeForce RTX 4060 Ti 8GB
- Best models for the NVIDIA GeForce RTX 3080 Ti
- Best models for the NVIDIA GeForce RTX 4090 Laptop GPU
Related Workloads
- Best local LLMs for cybersecurity
- Best local LLMs for document analysis
- Best local LLMs for fine-tuning
- Best local LLMs for ocr