Gemma 4 (Withdrawn Listings) — local AI model by Google DeepMind
Written by Jakub Rusinowski · Last updated
Google DeepMind's breakthrough open-weight model. Gemma 4 27B delivers GPT-4 level performance at 14 GB VRAM, making frontier-quality AI accessible on consumer GPUs. Features a hybrid architecture with interleaved local and global attention, multi-image understanding, and 128k context window. Achieves 85 tokens/second on RTX 4090.
Variants
The smallest Gemma 4 (Withdrawn Listings) variant needs about 3 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| Gemma 4 4B 4B | ~3.2 GB | 128,000 | ollama run gemma4:4b | |
| Gemma 4 12B 12B | ~8 GB | 128,000 | ollama run gemma-4-legacy | |
| Gemma 4 27B ⭐ 27B | ~17.1 GB | 128,000 | ollama run gemma4:27b |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run Gemma 4 (Withdrawn Listings) locally
Install Ollama, then pull the tag.
ollama run gemma4:4bPick a size above for its own VRAM figure, speed estimate and install command.
Licence
Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.
Applies to: Gemma 4 4B, Gemma 4 12B, Gemma 4 27B ⭐Recommended GPU
The cheapest catalogued GPU that runs Gemma 4 (Withdrawn Listings) locally (min 3 GB VRAM) is the Intel Arc B570 (10 GB).
Gemma 4 (Withdrawn Listings) — frequently asked questions
How much VRAM does Gemma 4 (Withdrawn Listings) need?
Gemma 4 (Withdrawn Listings) needs about 3 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Gemma 4 4B (3 GB, Q4_K_M); Gemma 4 12B (8 GB, Q4_K_M); Gemma 4 27B ⭐ (17 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Gemma 4 (Withdrawn Listings) on an RTX 4090 (24 GB)?
Yes — Gemma 4 (Withdrawn Listings) runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for Gemma 4 (Withdrawn Listings)?
Q4_K_M is the best balance of quality and VRAM for Gemma 4 (Withdrawn Listings) in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Gemma 4 (Withdrawn Listings) with Ollama?
Install Ollama, then run: ollama run gemma4:4b. This downloads Gemma 4 (Withdrawn Listings) and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.