Best Local LLMs for Mathematics

Written by Jakub Rusinowski · Last updated June 26, 2026

Symbolic and numerical problem solving, proofs, and step-by-step quantitative reasoning.

Top pick: Qwen 3 32B

Scores 100/100 for mathematics and quantitative work. 33B parameters, needing about 20.6 GB at Q4_K_M, 125K context, Apache 2.0.

Ranked for mathematics and quantitative work

ModelScoreParamsContextLicenceQuality index
1. Qwen 3 32B10033B125KApache 2.0— (estimated)
2. DeepSeek R1 Distill Qwen 32B98.832B128KMIT87 (cited)
3. GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning)98.232B125KApache-2.0— (estimated)
4. Qwen 3.7 35B-A3B96.135B256KApache-2.0— (estimated)
5. Gemma 4 31B9631B250KApache-2.0— (estimated)
6. Qwen 2.5 72B Instruct9672B128KApache-2.082 (cited)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirQwen 3 8B (93.9)
GLM-4.7 9B (92.8)
Qwen 3.5 7B (92.7)
12 GBRTX 3060 12 GB, RTX 5070Qwen 3 14B (99.9)
DeepSeek R1 Distill Qwen 14B (97.9)
Qwen 3.5 14B (94.8)
16 GBRTX 5080, RTX 4080, RX 9070 XTQwen 3 14B (99.9)
DeepSeek R1 Distill Qwen 14B (97.7)
Qwen 3.5 14B (94.6)
24 GBRTX 4090, RTX 3090, RX 7900 XTXDeepSeek R1 Distill Qwen 32B (100)
GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (100)
Qwen 3 32B (100)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBQwen 3 32B (100)
Qwen 2.5 72B Instruct (100)
DeepSeek R1 Distill Qwen 32B (99.3)
128 GB+Mac Studio, DGX Spark, multi-GPUQwen 3 32B (98.1)
Qwen 2.5 72B Instruct (98)
Qwen 3.5 122B-A10B (MoE) (97.7)

How this ranking works

The most concentrated capability weighting in the set: 85% reasoning, no creative weight at all. The quality floor of 70 is the highest anywhere here, because arithmetic and proof errors are silent — a plausible-looking wrong answer is worse than a refusal.

Worked example — Qwen 3 32B: capability 96.9 × 0.443, quality 93.7 × 0.261, context 100 × 0.169, license 70 × 0.029, accessibility 80 × 0.098 + 6 tag bonus (complex-reasoning, math).

Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 70, licence weight 0.25, latency weight 0.25.

Running mathematics and quantitative work locally

FAQ

What is the best local LLM for mathematics and quantitative work?

Qwen 3 32B, scoring 100/100 against this workload's published requirements. 106 models qualified.

What hardware do I need for mathematics and quantitative work?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

The most concentrated capability weighting in the set: 85% reasoning, no creative weight at all. The quality floor of 70 is the highest anywhere here, because arithmetic and proof errors are silent — a plausible-looking wrong answer is worse than a refusal.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware