Written by Jakub Rusinowski · Last updated July 21, 2026
Ranked for security analysis and defensive research: Log and alert triage, code review for vulnerabilities, and analysis that cannot be sent to a third party.
Best overall: AMD Ryzen AI Max+ 395
96 GB VRAM at 256 GB/s. It runs 90 of the models that qualify for this workload; the strongest is Qwen 3.7 35B-A3B at an estimated 60.3 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Ryzen AI Max+ 395 | 96 GB | $1,999 | 90 | Qwen 3.7 35B-A3B | ~60.3 tok/s |
| Best value | AMD Radeon RX 7900 XTX | 24 GB | $999 | 77 | Gemma 4 31B | ~33.9 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 53 | Qwen 3 14B | ~27.8 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 90 | Qwen 3.7 35B-A3B | ~63.8 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Ryzen AI Max+ 395 | 81.9 | 96 GB | $1,999 | 90 | ~60.3 | 0.5 | $22 |
| NVIDIA DGX Spark | 79.2 | 128 GB | $4,699 | 90 | ~63.8 | 0.43 | $52 |
| NVIDIA RTX 6000 Ada Generation | 76.3 | 48 GB | $6,799 | 82 | ~164.8 | 0.55 | $83 |
| NVIDIA L40S | 75.8 | 48 GB | $7,499 | 82 | ~154.1 | 0.44 | $91 |
| NVIDIA GeForce RTX 5090 | 73.6 | 32 GB | $1,999 | 77 | ~59.4 | 0.1 | $26 |
| NVIDIA GeForce RTX 4090 | 67.1 | 24 GB | $1,599 | 77 | ~35.5 | 0.08 | $21 |
| AMD Radeon RX 7900 XTX | 66.8 | 24 GB | $999 | 77 | ~33.9 | 0.1 | $13 |
| NVIDIA GeForce RTX 3090 | 63.8 | 24 GB | $1,499 | 77 | ~33.1 | 0.09 | $19 |
| AMD Radeon RX 7900 XT | 58.8 | 20 GB | $899 | 67 | ~28.6 | 0.09 | $13 |
| NVIDIA GeForce RTX 5060 | 47.9 | 8 GB | $299 | 43 | ~54.6 | 0.38 | $7 |
| AMD Radeon RX 9060 XT 8GB | 47.7 | 8 GB | $299 | 43 | ~40.4 | 0.27 | $7 |
| AMD Radeon RX 7800 XT | 46.7 | 16 GB | $499 | 61 | ~43.9 | 0.17 | $8 |
The AMD Ryzen AI Max+ 395 — 96 GB of VRAM runs 90 qualifying models, the strongest being Qwen 3.7.
The Intel Arc B570 at $219, which runs 53 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.