Gemma 3n E4B — 显存、速度与本地部署
作者: Jakub Rusinowski · 最后更新:
The full-power Gemma 3n variant. Processes text, images, audio, and video. Runs on devices with 4 GB RAM, including mid-range phones and all laptops.
Gemma 3n E4B 在 Q4_K_M 下约需 6 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。
按量化级别的显存与速度
计算基准:NVIDIA RTX 4090 (24 GB)。仅含权重与开销:该模型架构未公开,因此未计入 KV 缓存。
| 量化 | 显存 | 显存 | 速度(估算) | 适配 |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 3.4 GB | ~213 tok/s | 可运行 | |
| Q3_K_M 3.41 bpw | 4.1 GB | ~193 tok/s | 可运行 | |
| Q4_K_M 4.83 bpw | 5.5 GB | ~164 tok/s | 可运行 | |
| Q5_K_M 5.67 bpw | 6.4 GB | ~151 tok/s | 可运行 | |
| Q6_K 6.56 bpw | 7.2 GB | ~139 tok/s | 可运行 | |
| Q8_0 8.50 bpw | 9.1 GB | ~118 tok/s | 可运行 | |
| F16 16.00 bpw | 16.5 GB | ~75 tok/s | 可运行 |
黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Gemma 3n E4B 显存计算器 →
运行 Gemma 3n E4B
目录中能运行 Gemma 3n E4B 的最便宜 GPU 是 Intel Arc B570 (10 GB).
或在 Vast.ai 比较,低至 $0.35/小时 (typical low · varies)
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
如何运行 Gemma 3n E4B
安装 Ollama,然后运行:
ollama run gemma3n:e4b规格
Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.
- 参数量
- 4 Billion effective
- 上下文窗口
- 32,768
- 架构
- MatFormer, multimodal
- 提供商
- Google DeepMind
- 许可证
- Gemma ToS
- 规格量化
- Q4_K_M
- 系统内存
- 8 GB
- 记录更新于
- 2025-04-01
Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.
质量与使用场景
评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。
我的 GPU 能运行 Gemma 3n E4B 吗?
- Gemma 3n E4B 在 AMD Radeon RX 9060 XT 8GB 上
- Gemma 3n E4B 在 Intel Arc B570 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 3070 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 3070 Ti 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 3080 (10GB) 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 4060 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 5060 上
- Gemma 3n E4B 在 NVIDIA GeForce RTX 5060 Ti 8GB 上
Gemma 3n 的其他尺寸
Gemma 3n E4B — 常见问题
How much VRAM does Gemma 3n E4B need?
About 6 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Gemma 3n E4B run on an RTX 4090 (24 GB)?
Yes. Gemma 3n E4B needs about 6 GB at Q4_K_M, inside a 24 GB card, at an estimated 164 tokens/sec.
How do I run Gemma 3n E4B locally?
Install Ollama and run `ollama run gemma3n:e4b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does Gemma 3n come in?
Gemma 3n E2B (4 GB), Gemma 3n E4B (6 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.