MiniCPM-V 4.6 — VRAM, Speed & Local Setup

Written by Jakub Rusinowski · Last updated September 11, 2026

Model libraryMiniCPM-V → MiniCPM-V 4.6

1.3 billion parameters — about 1.6 GB at Q4_K_M including cache and overhead, which is genuinely phone-sized. A SigLIP2-400M vision encoder on a Qwen3.5-0.8B backbone, following LLaVA-UHD v4, with mixed 4x/16x visual token compression that cuts visual encoding FLOPs by more than 50%. Takes text, images and video (up to 128 frames) across a 262,144-token context. Apache 2.0, with day-one llama.cpp, Ollama, vLLM and SGLang support.

MiniCPM-V 4.6 needs about 2 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Specifications

Parameters1.3 Billion
Context window262,144
ArchitectureSigLIP2-400M encoder + Qwen3.5-0.8B decoder
ProviderOpenBMB
LicenceApache 2.0
Specified atQ4_K_M
System RAM8 GB
Record updated2026-09-11

Licence

Apache-2.0commercial use permitted. Commercial use permitted. No usage restrictions beyond attribution.

VRAM and Speed by Quantization

Modelled on a reference NVIDIA RTX 4090 (24 GB), with no KV cache (this record has no published architecture). Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.

QuantWeightsVRAM neededEst. speedFit on 24 GB
Q2_K0.4 GB1.2 GB~309 tok/s (est.)Fits comfortably
Q3_K_M0.6 GB1.4 GB~294 tok/s (est.)Fits comfortably
Q4_K_M0.8 GB1.6 GB~270 tok/s (est.)Fits comfortably
Q5_K_M0.9 GB1.7 GB~257 tok/s (est.)Fits comfortably
Q6_K1.1 GB1.9 GB~245 tok/s (est.)Fits comfortably
Q8_01.4 GB2.2 GB~222 tok/s (est.)Fits comfortably
F162.6 GB3.4 GB~164 tok/s (est.)Fits comfortably

Want the memory numbers alone, at every quantization level and your own context length? Use the MiniCPM-V 4.6 VRAM calculator.

Buy This HardwareIntel Arc B570 10GB — 10 GB VRAM · 150 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Recommended GPU

The cheapest catalogued GPU that runs MiniCPM-V 4.6 is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.
Check price on Amazon

How to Run MiniCPM-V 4.6

Install Ollama, then run:

ollama run minicpm-v4.6

Weights on Hugging Face: openbmb/MiniCPM-V-4.6.

Best for: multimodal, vision, edge devices, mobile.

Can I Run MiniCPM-V 4.6 on My GPU?

Other MiniCPM-V Sizes

MiniCPM-V 4.6 — Frequently Asked Questions

How much VRAM does MiniCPM-V 4.6 need?
About 2 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does MiniCPM-V 4.6 run on an RTX 4090 (24 GB)?
Yes. MiniCPM-V 4.6 needs about 2 GB at Q4_K_M, inside a 24 GB card, at an estimated 270 tokens/sec.
How do I run MiniCPM-V 4.6 locally?
Install Ollama and run `ollama run minicpm-v4.6`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does MiniCPM-V come in?
MiniCPM-V 4.6 (2 GB), MiniCPM-V 4.5 (6 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.

← All MiniCPM-V models | VRAM calculator | Check your own hardware