Written by Jakub Rusinowski · Last updated July 16, 2026
Reading papers and technical material, extracting arguments, and comparing sources.
Top pick: Gemma 4 31B
Scores 95.7/100 for research and technical reading. 31B parameters, needing about 19.5 GB at Q4_K_M, 250K context, Apache-2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Gemma 4 31B | 95.7 | 31B | 250K | Apache-2.0 | — (estimated) |
| 2. Inkling (BF16) | 95.5 | 1000B | 977K | Apache 2.0 | — (estimated) |
| 3. Qwen 3.6 27B | 94.8 | 28B | 256K | Apache-2.0 | — (estimated) |
| 4. DeepSeek V4-Pro | 94.2 | 1600B | 977K | MIT | — (estimated) |
| 5. DeepSeek V4.1 | 94.2 | 1600B | 977K | MIT | — (estimated) |
| 6. Qwen 3.7 35B-A3B | 93.4 | 35B | 256K | Apache-2.0 | — (estimated) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Qwen 3 8B (79.7) Qwen 3.5 9B (79.3) GLM-6 9B (79.3) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (82.7) Qwen 3.5 14B (80.9) DeepSeek R1 Distill Qwen 14B (80.5) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3 14B (82.7) Mistral Small 3.1 24B (82.2) Qwen 3.5 14B (80.7) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Gemma 4 31B (97.1) Qwen 3.6 27B (96.2) Qwen 3.7 35B-A3B (94.8) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Gemma 4 31B (96) Qwen 3.6 27B (94.8) Qwen 3.7 35B-A3B (94) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Gemma 4 31B (94.2) Qwen 3.6 27B (93.1) Llama 4.5 Scout (92.9) |
Reasoning-led (70%) with a 32K context floor, because a single paper with references routinely exceeds 30K tokens. Latency is down-weighted to 0.3 — research reading is not interactive in the way chat is.
Worked example — Gemma 4 31B: capability 93 × 0.454, quality 91.7 × 0.267, context 97.3 × 0.173, license 100 × 0.036, accessibility 80 × 0.07 + 3 tag bonus (long-context).
Requirements applied: context floor 32,768 tokens (ideal 262,144), quality floor 68, licence weight 0.3, latency weight 0.3.
Gemma 4 31B, scoring 95.7/100 against this workload's published requirements. 106 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Reasoning-led (70%) with a 32K context floor, because a single paper with references routinely exceeds 30K tokens. Latency is down-weighted to 0.3 — research reading is not interactive in the way chat is.