Best Local LLMs for Fine-tuning
Written by Jakub Rusinowski · Last updated June 26, 2026
Picking a base model to adapt on your own data with LoRA/QLoRA.
Top pick: Qwen 3.7 35B-A3B
Scores 92.8/100 for fine-tuning a base model. 35B parameters, needing about 21.9 GB at Q4_K_M, 256K context, Apache-2.0.
Ranked for fine-tuning a base model
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Qwen 3.7 35B-A3B | 92.8 | 35B | 256K | Apache-2.0 | — (estimated) |
| 2. Gemma 4 31B | 92.4 | 31B | 250K | Apache-2.0 | — (estimated) |
| 3. Qwen 3.6 35B-A3B | 92.2 | 35B | 256K | Apache-2.0 | — (estimated) |
| 4. DeepSeek R1 Distill Qwen 32B | 91.6 | 32B | 128K | MIT | 87 (cited) |
| 5. Qwen 3 14B | 91.2 | 15B | 125K | Apache 2.0 | — (estimated) |
| 6. GLM-6 9B | 91 | 9B | 125K | MIT | — (estimated) |
Best pick for your memory budget
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-6 9B (91) GLM-4 9B (90.6) Qwen3-Coder 8B (90) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (91.9) DeepSeek R1 Distill Qwen 14B (91.5) Gemma 4 12B (Unified) (90.2) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3 14B (91.9) Mistral Small 3.1 24B (91.4) DeepSeek R1 Distill Qwen 14B (91.3) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Qwen 3.7 35B-A3B (95.6) Gemma 4 31B (95.1) Qwen 3.6 35B-A3B (95) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Cogito v1 70B (96.4) GLM-5.1 72B (95.2) Nemotron Cascade 2 70B (94.7) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 3.5 122B-A10B (96) Cogito v1 70B (93.3) GPT-OSS 120B (93.3) |
How this ranking works
Ranked as a *base model* choice, not an assistant choice: licence weight is near-maximum (0.9) because a restrictive licence blocks distributing what you train, and latency is nearly ignored (0.2) because training throughput, not decode speed, sets the cost.
Worked example — Qwen 3.7 35B-A3B: capability 93 × 0.39, quality 92.7 × 0.23, context 100 × 0.149, license 100 × 0.093, accessibility 80 × 0.138.
Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 45, licence weight 0.9, latency weight 0.2.
Running fine-tuning a base model locally
- Fine-tuning memory is dominated by optimizer state and activations, not weights — see the fine-tuning VRAM calculator for the real figure.
- A smaller base model you can fully fine-tune usually beats a larger one you can only adapt at very low LoRA rank.
FAQ
What is the best local LLM for fine-tuning a base model?
Qwen 3.7 35B-A3B, scoring 92.8/100 against this workload's published requirements. 159 models qualified.
What hardware do I need for fine-tuning a base model?
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
How were these models ranked?
Ranked as a *base model* choice, not an assistant choice: licence weight is near-maximum (0.9) because a restrictive licence blocks distributing what you train, and latency is nearly ignored (0.2) because training throughput, not decode speed, sets the cost.
Hardware for This Workload
- Best GPU for fine-tuning
- Best models for the NVIDIA GeForce RTX 4060 Ti 8GB
- Best models for the NVIDIA GeForce RTX 3080 Ti
- Best models for the NVIDIA GeForce RTX 4090 Laptop GPU
Related Workloads
- Best local LLMs for enterprise assistant
- Best local LLMs for offline ai
- Best local LLMs for privacy
- Best local LLMs for document analysis