Written by Jakub Rusinowski · Last updated August 15, 2026
Work with data that must never leave your machine — legal, medical, financial or personal records.
Top pick: GPT-OSS 20B
Scores 93.4/100 for privacy-sensitive work. 20B parameters, needing about 12.9 GB at Q4_K_M, 128K context, Apache-2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. GPT-OSS 20B | 93.4 | 20B | 128K | Apache-2.0 | — (estimated) |
| 2. Qwen 3.7 35B-A3B | 92.5 | 35B | 256K | Apache-2.0 | — (estimated) |
| 3. Gemma 4 31B | 92.2 | 31B | 250K | Apache-2.0 | — (estimated) |
| 4. GLM-4.7-Flash 30B-A3B | 92.1 | 30B | 193K | MIT | — (estimated) |
| 5. Qwen 3.6 35B-A3B | 91.8 | 35B | 256K | Apache-2.0 | — (estimated) |
| 6. DeepSeek R1 Distill Qwen 32B | 91.1 | 32B | 128K | MIT | 87 (cited) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-6 9B (90.5) GLM-4.7 9B (90.4) GLM-5 9B (89.3) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3 14B (91.7) DeepSeek R1 Distill Qwen 14B (91.4) Qwen 2.5 14B Instruct (90.3) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | GPT-OSS 20B (95.2) Qwen 3 14B (91.7) Mistral Small 3.1 24B (91.4) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Qwen 3.7 35B-A3B (95.5) GLM-4.7-Flash 30B-A3B (95.2) Gemma 4 31B (95.2) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | GLM-5.1 72B (94.8) Qwen 2.5 72B Instruct (94) Nemotron 70B Instruct (93.8) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | GPT-oss 120B (99.1) Qwen 3.5 122B-A10B (93) Qwen 3.5 122B-A10B (MoE) (92.7) |
Weighted toward models that are practical to run entirely on hardware you own: a high licence weight (0.8) plus a moderate quality floor, rather than chasing the largest model. A model you cannot actually run locally scores nothing here no matter how capable it is.
Worked example — GPT-OSS 20B: capability 82.8 × 0.388, quality 82 × 0.228, context 100 × 0.148, license 100 × 0.082, accessibility 88 × 0.154 + 6 tag bonus (local-first, privacy-sensitive).
Requirements applied: context floor 16,384 tokens (ideal 131,072), quality floor 50, licence weight 0.8, latency weight 0.45.
GPT-OSS 20B, scoring 93.4/100 against this workload's published requirements. 111 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Weighted toward models that are practical to run entirely on hardware you own: a high licence weight (0.8) plus a moderate quality floor, rather than chasing the largest model. A model you cannot actually run locally scores nothing here no matter how capable it is.