MBZUAI Institute of Foundation Models36B (4B active)~23 GB VRAM at Q4_K_M

K2 Horizon MoVA 36B-A4B — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

A mixture-of-experts model with Mixture-of-Values attention: 37.4B parameters stored (the card rounds to 36B), about 4B active per token, so it generates quickly for its size. It scores 25 on the Artificial Analysis Intelligence Index, second only to Qwen3.8-27B among models that fit one consumer GPU when checked. Its card reports Terminal-Bench 2.1 58.6 against 44.9 for Qwen3.6-35B-A3B. The official Q4_K_M GGUF is 22.4 GB, and its plain grouped-query cache adds about 6.4 GB at 32K context, so this is a 32 GB-card model; a 24 GB card does not hold it at useful context. Apache 2.0, 512K native context.

K2 Horizon MoVA 36B-A4B needs about 23 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
14.7 GB~198 tok/sFits
Q3_K_M
3.41 bpw
18.4 GB~180 tok/sFits
Q4_K_M
4.83 bpw
25 GB~23 tok/sOffload
Q5_K_M
5.67 bpw
28.9 GB~21 tok/sOffload
Q6_K
6.56 bpw
33.1 GB~20 tok/sOffload
Q8_0
8.50 bpw
42.1 GB~17 tok/sOffload
F16
16.00 bpw
77.2 GB—Too big

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. K2 Horizon MoVA 36B-A4B VRAM calculator →

Get K2 Horizon MoVA 36B-A4B running

The cheapest catalogued GPU that runs K2 Horizon MoVA 36B-A4B is the AMD Radeon RX 7900 XTX (24 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 7900 XTX 24GB
24 GB VRAM · 355 W board power
2026 prices are volatile — check the current listing.

How to run K2 Horizon MoVA 36B-A4B

No Ollama library tag yet. Run the official GGUF with llama.cpp: the files declare the k2-horizon architecture, which mainline llama.cpp supports; an older Ollama or LM Studio build may not load it. The model card's own serving route is vLLM or SGLang.

Weights on Hugging Face: IFM/K2-Horizon-MoVA-36B-A4B ↗

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
37.4 Billion (4B active)
Context window
524,288
Architecture
Mixture-of-Experts (100 routed + 1 shared) with Mixture-of-Values attention
Provider
MBZUAI Institute of Foundation Models
Licence
Apache 2.0
Specified at
Q4_K_M
System RAM
64 GB
Record updated
2026-10-08
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forcodingagentic codingsoftware engineeringreasoninglong context
BenchmarkScoreProvenance
Terminal-Bench 2.158.6 / 100 %reported · https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B

Other K2 Horizon sizes

K2 Horizon MoVA 36B-A4B — frequently asked questions

How much VRAM does K2 Horizon MoVA 36B-A4B need?

About 23 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does K2 Horizon MoVA 36B-A4B run on an RTX 4090 (24 GB)?

Yes. K2 Horizon MoVA 36B-A4B needs about 23 GB at Q4_K_M, inside a 24 GB card, at an estimated 23 tokens/sec.

How do I run K2 Horizon MoVA 36B-A4B locally?

No Ollama library tag yet. Run the official GGUF with llama.cpp: the files declare the k2-horizon architecture, which mainline llama.cpp supports; an older Ollama or LM Studio build may not load it. The model card's own serving route is vLLM or SGLang. Running the published tag would send your prompts to a hosted GPU rather than your own machine.

What other sizes does K2 Horizon come in?

K2 Horizon 3.7B (4 GB), K2 Horizon 7B (6 GB), K2 Horizon MoVA 36B-A4B (23 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.