Written by Jakub Rusinowski · Last updated June 26, 2026
Symbolic and numerical problem solving, proofs, and step-by-step quantitative reasoning.
Top pick: Qwen 3 32B
Scores 100/100 for mathematics and quantitative work. 33B parameters, needing about 20.6 GB at Q4_K_M, 125K context, Apache 2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Qwen 3 32B | 100 | 33B | 125K | Apache 2.0 | — (estimated) |
| 2. DeepSeek R1 Distill Qwen 32B | 98.8 | 32B | 128K | MIT | 87 (cited) |
| 3. GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) | 98.2 | 32B | 125K | Apache-2.0 | — (estimated) |
| 4. Qwen 3.7 35B-A3B | 96.1 | 35B | 256K | Apache-2.0 | — (estimated) |
| 5. Gemma 4 31B | 96 | 31B | 250K | Apache-2.0 | — (estimated) |
| 6. Qwen 2.5 72B Instruct | 96 | 72B | 128K | Apache-2.0 | 82 (cited) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Qwen 3 8B (93.9) GLM-4.7 9B (92.8) Qwen 3.5 7B (92.7) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (99.9) DeepSeek R1 Distill Qwen 14B (97.9) Qwen 3.5 14B (94.8) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3 14B (99.9) DeepSeek R1 Distill Qwen 14B (97.7) Qwen 3.5 14B (94.6) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | DeepSeek R1 Distill Qwen 32B (100) GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (100) Qwen 3 32B (100) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Qwen 3 32B (100) Qwen 2.5 72B Instruct (100) DeepSeek R1 Distill Qwen 32B (99.3) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 3 32B (98.1) Qwen 2.5 72B Instruct (98) Qwen 3.5 122B-A10B (MoE) (97.7) |
The most concentrated capability weighting in the set: 85% reasoning, no creative weight at all. The quality floor of 70 is the highest anywhere here, because arithmetic and proof errors are silent — a plausible-looking wrong answer is worse than a refusal.
Worked example — Qwen 3 32B: capability 96.9 × 0.443, quality 93.7 × 0.261, context 100 × 0.169, license 70 × 0.029, accessibility 80 × 0.098 + 6 tag bonus (complex-reasoning, math).
Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 70, licence weight 0.25, latency weight 0.25.
Qwen 3 32B, scoring 100/100 against this workload's published requirements. 106 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
The most concentrated capability weighting in the set: 85% reasoning, no creative weight at all. The quality floor of 70 is the highest anywhere here, because arithmetic and proof errors are silent — a plausible-looking wrong answer is worse than a refusal.