Gemma 4 (Withdrawn Listings) — local AI model by Google DeepMind

Written by Jakub Rusinowski · Last updated

Withdrawn listing. This was an early, unverified listing. The verified data now lives on the Gemma 4 page. View Gemma 4 →

Google DeepMind's breakthrough open-weight model. Gemma 4 27B delivers GPT-4 level performance at 14 GB VRAM, making frontier-quality AI accessible on consumer GPUs. Features a hybrid architecture with interleaved local and global attention, multi-image understanding, and 128k context window. Achieves 85 tokens/second on RTX 4090.

Variants

The smallest Gemma 4 (Withdrawn Listings) variant needs about 3 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.

ModelVRAM
Gemma 4 4B
4B
~3.2 GB
Gemma 4 12B
12B
~8 GB
Gemma 4 27B ⭐
27B
~17.1 GB

Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.

How to run Gemma 4 (Withdrawn Listings) locally

Install Ollama, then pull the tag.

ollama run gemma4:4b

Pick a size above for its own VRAM figure, speed estimate and install command.

Licence

Gemma TermsCommercial use permitted

Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.

Applies to: Gemma 4 4B, Gemma 4 12B, Gemma 4 27B ⭐

Recommended GPU

The cheapest catalogued GPU that runs Gemma 4 (Withdrawn Listings) locally (min 3 GB VRAM) is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.

Gemma 4 (Withdrawn Listings) — frequently asked questions

How much VRAM does Gemma 4 (Withdrawn Listings) need?

Gemma 4 (Withdrawn Listings) needs about 3 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Gemma 4 4B (3 GB, Q4_K_M); Gemma 4 12B (8 GB, Q4_K_M); Gemma 4 27B ⭐ (17 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run Gemma 4 (Withdrawn Listings) on an RTX 4090 (24 GB)?

Yes — Gemma 4 (Withdrawn Listings) runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.

What quantization should I use for Gemma 4 (Withdrawn Listings)?

Q4_K_M is the best balance of quality and VRAM for Gemma 4 (Withdrawn Listings) in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Gemma 4 (Withdrawn Listings) with Ollama?

Install Ollama, then run: ollama run gemma4:4b. This downloads Gemma 4 (Withdrawn Listings) and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.