Best GPU for Local AI Agents
Written by Jakub Rusinowski · Last updated October 7, 2026
Ranked for autonomous agents and tool use: Multi-step autonomous loops where the model calls tools, reads results and decides what to do next.
Best overall: AMD Radeon RX 7900 XTX
24 GB VRAM at 960 GB/s. It runs 32 of the models that qualify for this workload; the strongest is Ternary Bonsai 27B at an estimated 38.3 tokens/sec.
The picks
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Radeon RX 7900 XTX | 24 GB | $999 | 32 | Ternary Bonsai 27B | ~38.3 tok/s |
| Best value | Intel Arc B570 | 10 GB | $219 | 10 | GLM-6 9B | ~42.7 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 10 | GLM-6 9B | ~42.7 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 40 | Devstral-2 123B | ~2.7 tok/s |
Full ranking for autonomous agents and tool use
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | 74.9 | 24 GB | $999 | 32 | ~38.3 | 0.11 | $31 |
| NVIDIA GeForce RTX 4090 | 73.9 | 24 GB | $1,599 | 32 | ~40.1 | 0.09 | $50 |
| NVIDIA GeForce RTX 3090 Ti | 73.3 | 24 GB | $1,999 | 32 | ~40.1 | 0.09 | $62 |
| NVIDIA GeForce RTX 5070 Ti | 73.1 | 16 GB | $749 | 15 | ~40.3 | 0.13 | $50 |
| NVIDIA GeForce RTX 5090 | 73 | 32 GB | $1,999 | 32 | ~66.6 | 0.12 | $62 |
| NVIDIA GeForce RTX 3090 | 72.4 | 24 GB | $1,499 | 32 | ~37.4 | 0.11 | $47 |
| NVIDIA GeForce RTX 5080 | 71 | 16 GB | $999 | 15 | ~42.9 | 0.12 | $67 |
| NVIDIA RTX 6000 Ada Generation | 70.8 | 48 GB | $6,799 | 32 | ~38.3 | 0.13 | $212 |
| AMD Radeon RX 7900 XT | 70.6 | 20 GB | $899 | 24 | ~32.4 | 0.1 | $37 |
| NVIDIA L40S | 67.4 | 48 GB | $7,499 | 32 | ~34.8 | 0.1 | $234 |
| AMD Radeon RX 9070 | 67.3 | 16 GB | $549 | 15 | ~29.6 | 0.13 | $37 |
| AMD Radeon RX 7800 XT | 67.2 | 16 GB | $499 | 15 | ~28.9 | 0.11 | $33 |
How these numbers are calculated
- GPUs are scored on four axes: capability (45%), speed (30%), value (20%) and efficiency (5%), each normalised across the whole ranking.
- Capability means the intrinsic strength of the best model the card can hold — not how many models fit, and not how fast it streams a small one.
- Speed is capped at 40 tok/s: past that, more throughput does not change how the model feels to use.
- Apple Silicon is excluded from this ranking. Those entries price a whole computer and rate a chip's package power, so on price-per-capability and performance-per-watt they would beat every add-in card by construction. Apple hardware is covered on the macOS platform page instead.
- Reliable structured tool calls are a hard gate here, not a bonus: a model with no tool-use tag is excluded regardless of score. Coding weight is high (40%) because tool arguments are effectively code. Context needs match RAG — an agent loop accumulates every observation it has ever seen.
FAQ
What is the best GPU for autonomous agents and tool use?
The AMD Radeon RX 7900 XTX — 24 GB of VRAM runs 32 qualifying models, the strongest being Bonsai 27B.
What is the cheapest GPU that works for autonomous agents and tool use?
The Intel Arc B570 at $219, which runs 10 qualifying models.
How much VRAM do I need for autonomous agents and tool use?
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.
What These GPUs Run
- Best models for the AMD Radeon RX 7900 XTX
- Best models for the Intel Arc B570
- Best models for the NVIDIA DGX Spark
GPU Reviews
- AMD Radeon RX 7900 XTX review
- NVIDIA GeForce RTX 4090 review
- NVIDIA GeForce RTX 3090 Ti review
- NVIDIA GeForce RTX 5070 Ti review