Written by Jakub Rusinowski · Last updated June 26, 2026
Translating between languages and working in a language other than English.
Top pick: GLM-4.7 9B
Scores 96.1/100 for translation and multilingual work. 9B parameters, needing about 6.2 GB at Q4_K_M, 125K context, Apache-2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. GLM-4.7 9B | 96.1 | 9B | 125K | Apache-2.0 | — (estimated) |
| 2. Aya Expanse 32B | 94 | 32B | 125K | CC-BY-NC | — (estimated) |
| 3. Qwen 3.6 27B | 93.6 | 28B | 256K | Apache-2.0 | — (estimated) |
| 4. Qwen 3.5 14B | 93.3 | 14B | 125K | Apache 2.0 | — (estimated) |
| 5. Qwen 3.7 35B-A3B | 92.1 | 35B | 256K | Apache-2.0 | — (estimated) |
| 6. Gemma 4 31B | 92 | 31B | 250K | Apache-2.0 | — (estimated) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-4.7 9B (96.1) Qwen 3.5 7B (92.1) GLM-6 9B (89.9) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | GLM-4.7 9B (95.1) Qwen 3.5 14B (94) Qwen 3 14B (92.1) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3.5 14B (93.8) GLM-4.7 9B (93.5) Mistral Small 3.1 24B (92.6) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Aya Expanse 32B (96.9) Qwen 3.6 27B (96.5) Qwen 3.7 35B-A3B (95) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Nemotron 70B Instruct (95.3) Aya Expanse 32B (94.7) GLM-5.1 72B (94.5) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 3.5 122B-A10B (MoE) (93.8) GPT-oss 120B (93) Qwen 3.5 122B-A10B (92.4) |
Split roughly evenly between reasoning and creative weight, since translation is simultaneously a fidelity and a fluency problem. Multilingual `bestFor` tagging carries more signal here than anywhere else, so the tag bonus frequently decides ranking between otherwise comparable models.
Worked example — GLM-4.7 9B: capability 85.4 × 0.414, quality 84.3 × 0.243, context 100 × 0.158, license 100 × 0.038, accessibility 100 × 0.146 + 6 tag bonus (multilingual, chinese).
Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 50, licence weight 0.35, latency weight 0.6.
GLM-4.7 9B, scoring 96.1/100 against this workload's published requirements. 111 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Split roughly evenly between reasoning and creative weight, since translation is simultaneously a fidelity and a fluency problem. Multilingual `bestFor` tagging carries more signal here than anywhere else, so the tag bonus frequently decides ranking between otherwise comparable models.