Best GPU for Local AI Fine-tuning
Written by Jakub Rusinowski · Last updated October 7, 2026
Ranked for fine-tuning a base model: Picking a base model to adapt on your own data with LoRA/QLoRA.
Best overall: Intel Arc Pro B70
32 GB VRAM at 608 GB/s. It runs 117 of the models that qualify for this workload; the strongest is Qwen 3.7 35B-A3B at an estimated 120.8 tokens/sec.
The picks
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | Intel Arc Pro B70 | 32 GB | $949 | 117 | Qwen 3.7 35B-A3B | ~120.8 tok/s |
| Best value | Intel Arc Pro B70 | 32 GB | $949 | 117 | Qwen 3.7 35B-A3B | ~120.8 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 81 | Qwen 3 14B | ~27.8 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 131 | Qwen 3.5 122B-A10B | ~25.8 tok/s |
Full ranking for fine-tuning a base model
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| Intel Arc Pro B70 | 82.5 | 32 GB | $949 | 117 | ~120.8 | 0.54 | $8 |
| AMD Radeon RX 7900 XTX | 81.1 | 24 GB | $999 | 117 | ~164.8 | 0.46 | $9 |
| AMD Radeon AI PRO R9700 | 80.4 | 32 GB | $1,299 | 117 | ~125.4 | 0.42 | $11 |
| NVIDIA GeForce RTX 3090 | 79.6 | 24 GB | $1,499 | 117 | ~162.2 | 0.46 | $13 |
| NVIDIA GeForce RTX 4090 | 78.9 | 24 GB | $1,599 | 117 | ~169.9 | 0.38 | $14 |
| NVIDIA GeForce RTX 3090 Ti | 78.3 | 24 GB | $1,999 | 117 | ~169.9 | 0.38 | $17 |
| NVIDIA GeForce RTX 5090 | 78 | 32 GB | $1,999 | 117 | ~233 | 0.41 | $17 |
| AMD Radeon RX 7900 XT | 63.7 | 20 GB | $899 | 106 | ~28.6 | 0.09 | $8 |
| Intel Arc B580 | 52.1 | 12 GB | $249 | 82 | ~33 | 0.17 | $3 |
| AMD Radeon RX 7800 XT | 50.3 | 16 GB | $499 | 92 | ~43.9 | 0.17 | $5 |
| AMD Radeon RX 9070 | 50 | 16 GB | $549 | 92 | ~44.9 | 0.2 | $6 |
| NVIDIA GeForce RTX 5070 | 49.6 | 12 GB | $549 | 82 | ~46.9 | 0.19 | $7 |
How these numbers are calculated
- GPUs are scored on four axes: capability (45%), speed (30%), value (20%) and efficiency (5%), each normalised across the whole ranking.
- Capability means the intrinsic strength of the best model the card can hold — not how many models fit, and not how fast it streams a small one.
- Speed is capped at 40 tok/s: past that, more throughput does not change how the model feels to use.
- Apple Silicon is excluded from this ranking. Those entries price a whole computer and rate a chip's package power, so on price-per-capability and performance-per-watt they would beat every add-in card by construction. Apple hardware is covered on the macOS platform page instead.
- Ranked as a *base model* choice, not an assistant choice: licence weight is near-maximum (0.9) because a restrictive licence blocks distributing what you train, and latency is nearly ignored (0.2) because training throughput, not decode speed, sets the cost.
FAQ
What is the best GPU for fine-tuning a base model?
The Intel Arc Pro B70 — 32 GB of VRAM runs 117 qualifying models, the strongest being Qwen 3.7.
What is the cheapest GPU that works for fine-tuning a base model?
The Intel Arc B570 at $219, which runs 81 qualifying models.
How much VRAM do I need for fine-tuning a base model?
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.
What These GPUs Run
- Best models for the Intel Arc Pro B70
- Best models for the Intel Arc B570
- Best models for the NVIDIA DGX Spark
GPU Reviews
- Intel Arc Pro B70 review
- AMD Radeon RX 7900 XTX review
- AMD Radeon AI PRO R9700 review
- NVIDIA GeForce RTX 3090 review