Written by Jakub Rusinowski · Last updated July 15, 2026
A model that works with no internet connection at all — on a laptop, in the field, or inside an air-gapped network.
Top pick: Ternary Bonsai 27B
Scores 93.5/100 for fully offline and air-gapped use. 27B parameters, needing about 17.1 GB at Q4_K_M, 256K context, Apache 2.0.
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Ternary Bonsai 27B | 93.5 | 27B | 256K | Apache 2.0 | — (estimated) |
| 2. Gemma 4 12B (Unified) | 92.4 | 12B | 125K | Apache-2.0 | — (estimated) |
| 3. Qwen 3.7 35B-A3B | 92.1 | 35B | 256K | Apache-2.0 | — (estimated) |
| 4. Qwen 3 14B | 91.9 | 15B | 125K | Apache 2.0 | — (estimated) |
| 5. Qwen 3.6 35B-A3B | 91.5 | 35B | 256K | Apache-2.0 | — (estimated) |
| 6. Phi-4 Mini (3.8B) | 91.4 | 4B | 125K | MIT | — (estimated) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | GLM-6 9B (90.9) GLM-4.7 9B (90.6) Qwen 3 8B (90.1) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Gemma 4 12B (Unified) (93.2) Qwen 3 14B (92.7) DeepSeek R1 Distill Qwen 14B (91.3) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Qwen 3 14B (92.7) Mistral Small 3.1 24B (92.4) Gemma 4 12B (Unified) (91.8) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Ternary Bonsai 27B (96.9) Qwen 3.7 35B-A3B (95.6) Gemma 4 31B (95.1) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | GLM-5.1 72B (95.2) Nemotron 70B Instruct (94.7) Qwen 2.5 72B Instruct (94.1) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 3.5 122B-A10B (MoE) (93.7) GPT-oss 120B (93.4) Qwen 3.5 122B-A10B (93) |
The lowest quality floor in the set (40) and the strongest tag weighting toward small, low-memory models. Offline use is constrained by the hardware you happen to have with you, so a model that runs on 8 GB and never stalls beats a stronger model that needs a desktop.
Worked example — Ternary Bonsai 27B: capability 88.6 × 0.389, quality 87.7 × 0.229, context 100 × 0.149, license 70 × 0.062, accessibility 80 × 0.172 + 6 tag bonus (laptop, on-device).
Requirements applied: context floor 8,192 tokens (ideal 65,536), quality floor 40, licence weight 0.6, latency weight 0.7.
Ternary Bonsai 27B, scoring 93.5/100 against this workload's published requirements. 111 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
The lowest quality floor in the set (40) and the strongest tag weighting toward small, low-memory models. Offline use is constrained by the hardware you happen to have with you, so a model that runs on 8 GB and never stalls beats a stronger model that needs a desktop.