Best Local LLMs for Summarization
Written by Jakub Rusinowski · Last updated September 6, 2026
Condensing long inputs — transcripts, threads, reports — into accurate short output.
Top pick: Gemma 4 31B
Scores 94.6/100 for summarization. 31B parameters, needing about 19.5 GB at Q4_K_M, 250K context, Apache-2.0.
Ranked for summarization
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Gemma 4 31B | 94.6 | 31B | 250K | Apache-2.0 | — (estimated) |
| 2. Qwen 3.6 27B | 93.3 | 28B | 256K | Apache-2.0 | — (estimated) |
| 3. Qwen 3.5 14B | 92.3 | 14B | 125K | Apache 2.0 | — (estimated) |
| 4. Qwen3.8 27B | 92.3 | 28B | 256K | Apache-2.0 | — (estimated) |
| 5. Qwen 3.5 32B | 92.3 | 32B | 125K | Apache 2.0 | — (estimated) |
| 6. Gemma 4 12B | 92 | 12B | 125K | Gemma License (commercial OK) | — (estimated) |
Best pick for your memory budget
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-4 9B (90.1) Qwen 3 8B (89.8) GLM-6 9B (89.8) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3.5 14B (93.2) Gemma 4 12B (92.8) Qwen 3 14B (92.4) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3.5 14B (92.8) Mistral Small 3.1 24B (92.6) Qwen 3 14B (92.4) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Gemma 4 31B (98.1) Qwen 3.6 27B (96.8) Qwen3.8 27B (95.8) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | GLM-5.1 72B (97.4) Gemma 4 31B (95.2) Llama 3.3 70B Instruct (94.1) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 3.5 122B-A10B (95.6) Qwen3.8-Flash-Next (94) Qwen 3.5 122B-A10B (MoE) (93.8) |
How this ranking works
Reasoning-led but with a real creative weight (35%), because a summary is judged on readability as well as coverage. The low quality floor (45) reflects that summarization is the workload small models handle best relative to their size.
Worked example — Gemma 4 31B: capability 92.8 × 0.401, quality 91.7 × 0.236, context 100 × 0.153, license 100 × 0.032, accessibility 80 × 0.177 + 3 tag bonus (long-context).
Requirements applied: context floor 32,768 tokens (ideal 131,072), quality floor 40, licence weight 0.3, latency weight 0.5.
Running summarization locally
- Summarization is dominated by prompt processing, not generation: the input is long and the output is short, so prompt tokens/sec is the number to watch.
- This is the strongest case for a small model — a 7–8B model summarises far closer to a frontier model than it reasons.
FAQ
What is the best local LLM for summarization?
Gemma 4 31B, scoring 94.6/100 against this workload's published requirements. 159 models qualified.
What hardware do I need for summarization?
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
How were these models ranked?
Reasoning-led but with a real creative weight (35%), because a summary is judged on readability as well as coverage. The low quality floor (45) reflects that summarization is the workload small models handle best relative to their size.
Hardware for This Workload
- Best GPU for summarization
- Best models for the NVIDIA GeForce RTX 4060 Ti 8GB
- Best models for the NVIDIA GeForce RTX 3080 Ti
- Best models for the NVIDIA GeForce RTX 4090 Laptop GPU
Related Workloads
- Best local LLMs for translation
- Best local LLMs for vision
- Best local LLMs for rag
- Best local LLMs for privacy