K2 Horizon — local AI model by MBZUAI Institute of Foundation Models
Written by Jakub Rusinowski · Last updated
A fully open model family from MBZUAI's Institute of Foundation Models — weights, training data, recipe and code all published under Apache 2.0, with intermediate checkpoints released too. Three sizes: a 3.7B and a 7B dense model for small machines, and a 36B mixture-of-experts model that activates about 4B parameters per token. All three have a 512K native context. On the independent Artificial Analysis Intelligence Index they score 16, 21 and 25, the strongest results in their size classes when they were checked.
Variants
The smallest K2 Horizon variant needs about 4 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| K2 Horizon 3.7B → 3.7B | ~3.9 GB | 524,288 | cloud-hosted tag | |
| K2 Horizon 7B → 7B | ~6.2 GB | 524,288 | cloud-hosted tag | |
| K2 Horizon MoVA 36B-A4B → 36B (4B active) | ~23.4 GB | 524,288 | cloud-hosted tag |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run K2 Horizon locally
Install Ollama, then pull the tag.
No Ollama library tag yet. Run the official GGUF with llama.cpp: the files declare the k2-horizon architecture, which mainline llama.cpp supports; an older Ollama or LM Studio build may not load it. The model card's own serving route is vLLM or SGLang.
Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: K2 Horizon 3.7B, K2 Horizon 7B, K2 Horizon MoVA 36B-A4BRecommended GPU
The cheapest catalogued GPU that runs K2 Horizon locally (min 4 GB VRAM) is the Intel Arc B570 (10 GB).
K2 Horizon — frequently asked questions
How much VRAM does K2 Horizon need?
K2 Horizon needs about 4 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: K2 Horizon 3.7B (4 GB, Q4_K_M); K2 Horizon 7B (6 GB, Q4_K_M); K2 Horizon MoVA 36B-A4B (23 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run K2 Horizon on an RTX 4090 (24 GB)?
Yes — K2 Horizon runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for K2 Horizon?
Q4_K_M is the best balance of quality and VRAM for K2 Horizon in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run K2 Horizon with Ollama?
K2 Horizon has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.