Best Local LLMs for Agents

Written by Jakub Rusinowski · Last updated October 6, 2026

Multi-step autonomous loops where the model calls tools, reads results and decides what to do next.

Top pick: Ternary Bonsai 27B

Scores 91.7/100 for autonomous agents and tool use. 27B parameters, needing about 17.1 GB at Q4_K_M, 256K context, Apache 2.0.

Ranked for autonomous agents and tool use

ModelScoreParamsContextLicenceQuality index
1. Ternary Bonsai 27B91.727B256KApache 2.0— (estimated)
2. Devstral Small 2 24B91.224B256KApache-2.0— (estimated)
3. Nemotron-Cascade 2 30B-A3B90.632B1MNVIDIA Open Model License— (estimated)
4. Nemotron 3.5 Lightning 30B-A3B90.330B977KOpenMDW-1.1— (estimated)
5. Qwen3-Coder 30B-A3B (MoE)89.931B256KApache-2.0— (estimated)
6. Seed-OSS-36B-Instruct89.536B512KApache-2.0— (estimated)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirGLM-6 9B (86.8)
GLM-5 9B (85.4)
K2 Horizon 7B (78.6)
12 GBRTX 3060 12 GB, RTX 5070GLM-6 9B (86)
GLM-5 9B (84.6)
Qwen 3.5 14B (84.6)
16 GBRTX 5080, RTX 4080, RX 9070 XTDevstral Small 2 24B (92.6)
GLM-6 9B (84.7)
Qwen 3.5 14B (84.4)
24 GBRTX 4090, RTX 3090, RX 7900 XTXTernary Bonsai 27B (94)
Nemotron-Cascade 2 30B-A3B (92.8)
Devstral Small 2 24B (92.6)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBTernary Bonsai 27B (91.6)
Nemotron-Cascade 2 30B-A3B (91)
Nemotron 3.5 Lightning 30B-A3B (90.6)
128 GB+Mac Studio, DGX Spark, multi-GPUDevstral-2 123B (95.2)
Qwen3.8-Flash-Next (93.9)
Qwen3-Coder-Next (80B-A3B MoE) (93.1)

How this ranking works

Reliable structured tool calls are a hard gate here, not a bonus: a model with no tool-use tag is excluded regardless of score. Coding weight is high (40%) because tool arguments are effectively code. Context needs match RAG — an agent loop accumulates every observation it has ever seen.

Worked example — Ternary Bonsai 27B: capability 89.3 × 0.427, quality 87.7 × 0.251, context 100 × 0.163, license 70 × 0.045, accessibility 80 × 0.113 + 3 tag bonus (agents).

Requirements applied: context floor 32,768 tokens (ideal 262,144), quality floor 60, licence weight 0.4, latency weight 0.6. Tool-calling capability is required.

Running autonomous agents and tool use locally

FAQ

What is the best local LLM for autonomous agents and tool use?

Ternary Bonsai 27B, scoring 91.7/100 against this workload's published requirements. 55 models qualified.

What hardware do I need for autonomous agents and tool use?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

Reliable structured tool calls are a hard gate here, not a bonus: a model with no tool-use tag is excluded regardless of score. Coding weight is high (40%) because tool arguments are effectively code. Context needs match RAG — an agent loop accumulates every observation it has ever seen.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware