Best Local LLMs for Fine-tuning

Written by Jakub Rusinowski · Last updated June 26, 2026

Picking a base model to adapt on your own data with LoRA/QLoRA.

Top pick: Qwen 3.7 35B-A3B

Scores 92.8/100 for fine-tuning a base model. 35B parameters, needing about 21.9 GB at Q4_K_M, 256K context, Apache-2.0.

Ranked for fine-tuning a base model

ModelScoreParamsContextLicenceQuality index
1. Qwen 3.7 35B-A3B92.835B256KApache-2.0— (estimated)
2. Gemma 4 31B92.431B250KApache-2.0— (estimated)
3. Qwen 3.6 35B-A3B92.235B256KApache-2.0— (estimated)
4. DeepSeek R1 Distill Qwen 32B91.632B128KMIT87 (cited)
5. Qwen 3 14B91.215B125KApache 2.0— (estimated)
6. GLM-6 9B919B125KMIT— (estimated)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirGLM-6 9B (91)
GLM-4 9B (90.6)
Qwen3-Coder 8B (90)
12 GBRTX 3060 12 GB, RTX 5070Qwen 3 14B (91.9)
DeepSeek R1 Distill Qwen 14B (91.5)
Gemma 4 12B (Unified) (90.2)
16 GBRTX 5080, RTX 4080, RX 9070 XTQwen 3 14B (91.9)
Mistral Small 3.1 24B (91.4)
DeepSeek R1 Distill Qwen 14B (91.3)
24 GBRTX 4090, RTX 3090, RX 7900 XTXQwen 3.7 35B-A3B (95.6)
Gemma 4 31B (95.1)
Qwen 3.6 35B-A3B (95)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBCogito v1 70B (96.4)
GLM-5.1 72B (95.2)
Nemotron Cascade 2 70B (94.7)
128 GB+Mac Studio, DGX Spark, multi-GPUQwen 3.5 122B-A10B (96)
Cogito v1 70B (93.3)
GPT-OSS 120B (93.3)

How this ranking works

Ranked as a *base model* choice, not an assistant choice: licence weight is near-maximum (0.9) because a restrictive licence blocks distributing what you train, and latency is nearly ignored (0.2) because training throughput, not decode speed, sets the cost.

Worked example — Qwen 3.7 35B-A3B: capability 93 × 0.39, quality 92.7 × 0.23, context 100 × 0.149, license 100 × 0.093, accessibility 80 × 0.138.

Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 45, licence weight 0.9, latency weight 0.2.

Running fine-tuning a base model locally

FAQ

What is the best local LLM for fine-tuning a base model?

Qwen 3.7 35B-A3B, scoring 92.8/100 against this workload's published requirements. 159 models qualified.

What hardware do I need for fine-tuning a base model?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

Ranked as a *base model* choice, not an assistant choice: licence weight is near-maximum (0.9) because a restrictive licence blocks distributing what you train, and latency is nearly ignored (0.2) because training throughput, not decode speed, sets the cost.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware