Written by Jakub Rusinowski · Last updated July 21, 2026
Ranked for autonomous agents and tool use: Multi-step autonomous loops where the model calls tools, reads results and decides what to do next.
Best overall: AMD Radeon RX 7900 XTX
24 GB VRAM at 960 GB/s. It runs 17 of the models that qualify for this workload; the strongest is Ternary Bonsai 27B at an estimated 38.3 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Radeon RX 7900 XTX | 24 GB | $999 | 17 | Ternary Bonsai 27B | ~38.3 tok/s |
| Best value | AMD Radeon RX 7900 XT | 20 GB | $899 | 9 | Ternary Bonsai 27B | ~32.4 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 4 | GLM-6 9B | ~42.7 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 24 | Ternary Bonsai 27B | ~11.6 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | 78.7 | 24 GB | $999 | 17 | ~38.3 | 0.11 | $59 |
| NVIDIA GeForce RTX 4090 | 78 | 24 GB | $1,599 | 17 | ~40.1 | 0.09 | $94 |
| NVIDIA GeForce RTX 5090 | 76.9 | 32 GB | $1,999 | 17 | ~66.6 | 0.12 | $118 |
| NVIDIA GeForce RTX 3090 | 76 | 24 GB | $1,499 | 17 | ~37.4 | 0.11 | $88 |
| NVIDIA RTX 6000 Ada Generation | 74.7 | 48 GB | $6,799 | 18 | ~38.3 | 0.13 | $378 |
| AMD Radeon RX 7900 XT | 73.1 | 20 GB | $899 | 9 | ~32.4 | 0.1 | $100 |
| NVIDIA L40S | 70.4 | 48 GB | $7,499 | 18 | ~34.8 | 0.1 | $417 |
| Intel Arc B570 | 54.8 | 10 GB | $219 | 4 | ~42.7 | 0.28 | $55 |
| Intel Arc B580 | 50.9 | 12 GB | $249 | 5 | ~50.3 | 0.26 | $50 |
| NVIDIA GeForce RTX 5060 | 49.5 | 8 GB | $299 | 3 | ~49.5 | 0.34 | $100 |
| AMD Ryzen AI Max+ 395 | 47.4 | 96 GB | $1,999 | 24 | ~10.9 | 0.09 | $83 |
| NVIDIA GeForce RTX 3060 (12GB) | 47.1 | 12 GB | $329 | 5 | ~40.6 | 0.24 | $66 |
The AMD Radeon RX 7900 XTX — 24 GB of VRAM runs 17 qualifying models, the strongest being Bonsai 27B.
The Intel Arc B570 at $219, which runs 4 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.