Written by Jakub Rusinowski · Last updated June 26, 2026
Everyday questions, drafting, explanation and light analysis — the local replacement for a hosted chat assistant.
Top pick: Qwen 3 8B
Scores 96.1/100 for a general-purpose local assistant. 8B parameters, needing about 5.8 GB at Q4_K_M, 125K context, Apache 2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Qwen 3 8B | 96.1 | 8B | 125K | Apache 2.0 | — (estimated) |
| 2. Gemma 4 27B ⭐ | 94.2 | 27B | 125K | Gemma License (commercial OK) | — (estimated) |
| 3. Mistral Small 3.1 24B | 93.9 | 24B | 125K | Apache 2.0 | — (estimated) |
| 4. Qwen 3.5 14B | 93.6 | 14B | 125K | Apache 2.0 | — (estimated) |
| 5. GLM-6 9B | 93.3 | 9B | 125K | MIT | — (estimated) |
| 6. GLM-4.7 9B | 93.1 | 9B | 125K | Apache-2.0 | — (estimated) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Qwen 3 8B (96.1) GLM-6 9B (93.3) GLM-4.7 9B (93.1) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 8B (94.5) Qwen 3.5 14B (94.4) Gemma 3 12B Instruct (93.6) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Mistral Small 3.1 24B (95.9) Qwen 3.5 14B (94.1) Qwen 3.5 14B (93.2) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Gemma 4 27B ⭐ (97.4) Mistral Small 3.1 24B (95.9) Qwen 3.7 35B-A3B (95.2) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Nemotron 70B Instruct (98.3) GLM-5.1 72B (94.7) Qwen 3.5 72B (94.4) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | GPT-oss 120B (96) Nemotron 70B Instruct (94.7) Qwen 3.5 122B-A10B (MoE) (94.2) |
Balanced capability weighting with reasoning slightly ahead, since a general assistant is judged mostly on whether its answers hold up. Latency carries the highest weight of any workload (0.8): conversational use is the one case where waiting is immediately obvious. Context requirements are modest — most chat turns are short.
Worked example — Qwen 3 8B: capability 86.5 × 0.409, quality 86 × 0.24, context 100 × 0.156, license 70 × 0.032, accessibility 100 × 0.162 + 6 tag bonus (chat, balanced).
Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 50, licence weight 0.3, latency weight 0.8.
Qwen 3 8B, scoring 96.1/100 against this workload's published requirements. 111 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Balanced capability weighting with reasoning slightly ahead, since a general assistant is judged mostly on whether its answers hold up. Latency carries the highest weight of any workload (0.8): conversational use is the one case where waiting is immediately obvious. Context requirements are modest — most chat turns are short.