LLM Configurator is the definitive free tool for checking GPU compatibility with local LLMs. Enter your GPU's VRAM and system RAM to instantly discover which open-source AI models you can run — with Ollama install commands, speed estimates, and electricity cost calculations.
| VRAM | Models You Can Run |
|---|---|
| 2–4 GB | SmolLM2 1.7B, Phi-3.5 Mini, Gemma 4 E2B, Granite 4.1 3B |
| 6–8 GB | Llama 3.1 8B, Gemma 4 E4B, Phi-4 Mini, DeepSeek R1 8B, Granite 4.1 8B |
| 8–12 GB | Phi-4 14B (Q4), Qwen 2.5 14B (Q4), Mistral NeMo 12B, Gemma 4 E4B (FP16) |
| 12–16 GB | Llama 4 Scout 17B (Q4), Qwen 3.5 14B, Granite 4.1 30B (Q4) |
| 16–24 GB | Gemma 4 12B Unified, Qwen 3.5 27B, Mistral Small 4 (Q4), Qwen 2.5 Coder 32B |
| 24+ GB | Llama 3.3 70B (Q4), Llama 4 Maverick (Q4), DeepSeek R1 32B, Qwen 3.5 35B-A3B MoE |
Scout: 17B active / 109B total (MoE). Requires ~10 GB VRAM at Q4. ollama run llama4:scout. Maverick: 17B active / 400B total. Requires ~24 GB VRAM. ollama run llama4:maverick
Latest-generation DeepSeek flagship. Distilled and MoE variants span consumer to datacenter hardware — see the model page for exact VRAM per variant.
Google's 2026 open model family with compact E2B/E4B variants for low-VRAM machines and larger 12B-class models. ollama run gemma4
Alibaba's 2026 release with dense 14B/27B options plus efficient 35B-A3B MoE variants that fit modest GPUs. ollama run qwen3.5
24B multimodal model. ~24 GB VRAM at Q4. ollama run mistral-small4
LLM Configurator is a free, independent tool for the local AI community, created by Jakub Rusinowski, an AI educator and workshop leader on local LLM deployment. Supports 75+ open-source models. No account required. No ads. Free forever. Contact: contact@llmconfigurator.com