Best Local LLMs for Writing
Written by Jakub Rusinowski · Last updated June 26, 2026
Drafting, rewriting and editing prose where voice and readability matter as much as correctness.
Top pick: Gemma 3 27B Instruct
Scores 94.2/100 for writing and editing. 27B parameters, needing about 17.1 GB at Q4_K_M, 128K context, Gemma.
Ranked for writing and editing
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Gemma 3 27B Instruct | 94.2 | 27B | 128K | Gemma | 78 (cited) |
| 2. Gemma 4 27B ⭐ | 93.3 | 27B | 125K | Gemma License (commercial OK) | — (estimated) |
| 3. Mistral Small 3.1 24B | 93.2 | 24B | 125K | Apache 2.0 | — (estimated) |
| 4. Qwen 3.5 14B | 92.6 | 14B | 125K | Apache 2.0 | — (estimated) |
| 5. GLM-4 9B | 92.5 | 9B | 125K | Apache-2.0 | — (estimated) |
| 6. GLM-6 9B | 92.2 | 9B | 125K | MIT | — (estimated) |
Best pick for your memory budget
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-4 9B (92.5) GLM-6 9B (92.2) Qwen 3.5 7B (91.4) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Gemma 3 12B Instruct (94.3) Qwen 3.5 14B (93.4) Qwen 3.5 14B (92.6) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Mistral Small 3.1 24B (95.1) Qwen 3.5 14B (93.1) Gemma 3 12B Instruct (93) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Gemma 3 27B Instruct (97.3) Gemma 4 27B ⭐ (96.4) Mistral Small 3.1 24B (95.1) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Llama 3.3 70B Instruct (96.7) Gemma 3 27B Instruct (94.1) GLM-5.1 72B (93.8) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Llama 3.3 70B Instruct (93.2) Qwen 3.5 122B-A10B (MoE) (93.2) GPT-OSS 120B (92.6) |
How this ranking works
The only workload where the creative score dominates (65%). The quality floor is deliberately the lowest of any profile (45): prose quality correlates weakly with reasoning benchmarks, so excluding models on an aggregate index would discard genuinely good writing models.
Worked example — Gemma 3 27B Instruct: capability 94.2 × 0.412, quality 78 × 0.243, context 100 × 0.158, license 70 × 0.033, accessibility 80 × 0.155 + 6 tag bonus (chat, creative).
Requirements applied: context floor 8,192 tokens (ideal 131,072), quality floor 45, licence weight 0.3, latency weight 0.6.
Running writing and editing locally
- Long-form writing is where a large context window pays off in a way chat does not: the model has to hold the whole piece to keep its voice consistent.
- Heavier quantization degrades style before it degrades facts, so prefer Q5 or higher for prose if memory allows.
FAQ
What is the best local LLM for writing and editing?
Gemma 3 27B Instruct, scoring 94.2/100 against this workload's published requirements. 159 models qualified.
What hardware do I need for writing and editing?
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
How were these models ranked?
The only workload where the creative score dominates (65%). The quality floor is deliberately the lowest of any profile (45): prose quality correlates weakly with reasoning benchmarks, so excluding models on an aggregate index would discard genuinely good writing models.
Hardware for This Workload
- Best GPU for writing
- Best models for the NVIDIA GeForce RTX 4060 Ti 8GB
- Best models for the NVIDIA GeForce RTX 3080 Ti
- Best models for the NVIDIA GeForce RTX 4090 Laptop GPU
Related Workloads
- Best local LLMs for translation
- Best local LLMs for summarization
- Best local LLMs for general assistant
- Best local LLMs for vision