K2 Horizon 3.7B — VRAM, speed & local setup
Written by Jakub Rusinowski · Last updated
The small dense K2 Horizon: a 3.7B core plus a large 250K-token vocabulary, 5.1B parameters stored in all, so the official Q4_K_M GGUF is 3.2 GB. It scores 16 on the Artificial Analysis Intelligence Index. Its makers report 68.6 on SWE-bench Verified against 41.2 for Qwen3.5-4B, and level with it on Terminal-Bench 2.1 (25.1 against 25.8). The KV cache is a plain grouped-query cache, so long context is expensive on a small card: an 8 GB GPU holds it comfortably at 16K, not at 32K. Apache 2.0, 512K native context.
K2 Horizon 3.7B needs about 4 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
VRAM and speed by quantization
Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.
| Quant | Memory | VRAM | Speed (est.) | Fit |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 3.7 GB | ~190 tok/s | Fits | |
| Q3_K_M 3.41 bpw | 4.2 GB | ~169 tok/s | Fits | |
| Q4_K_M 4.83 bpw | 5.1 GB | ~140 tok/s | Fits | |
| Q5_K_M 5.67 bpw | 5.6 GB | ~128 tok/s | Fits | |
| Q6_K 6.56 bpw | 6.2 GB | ~117 tok/s | Fits | |
| Q8_0 8.50 bpw | 7.4 GB | ~98 tok/s | Fits | |
| F16 16.00 bpw | 12.2 GB | ~60 tok/s | Fits |
Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. K2 Horizon 3.7B VRAM calculator →
Get K2 Horizon 3.7B running
The cheapest catalogued GPU that runs K2 Horizon 3.7B is the Intel Arc B570 (10 GB).
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
How to run K2 Horizon 3.7B
No Ollama library tag yet. Run the official GGUF with llama.cpp: the files declare the k2-horizon architecture, which mainline llama.cpp supports; an older Ollama or LM Studio build may not load it. The model card's own serving route is vLLM or SGLang.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.
- Parameters
- 5.1 Billion (3.7B core + embeddings)
- Context window
- 524,288
- Architecture
- Dense Transformer (GQA)
- Provider
- MBZUAI Institute of Foundation Models
- Licence
- Apache 2.0
- Specified at
- Q4_K_M
- System RAM
- 16 GB
- Record updated
- 2026-10-08
Commercial use permitted. No usage restrictions beyond attribution.
Quality and use cases
Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.
| Benchmark | Score | Provenance |
|---|---|---|
| SWE-bench Verified | 68.6 / 100 % | reported · https://huggingface.co/IFM/K2-Horizon-3.7B |
| Terminal-Bench 2.1 | 25.1 / 100 % | reported · https://huggingface.co/IFM/K2-Horizon-3.7B |
Other K2 Horizon sizes
K2 Horizon 3.7B — frequently asked questions
How much VRAM does K2 Horizon 3.7B need?
About 4 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does K2 Horizon 3.7B run on an RTX 4090 (24 GB)?
Yes. K2 Horizon 3.7B needs about 4 GB at Q4_K_M, inside a 24 GB card, at an estimated 140 tokens/sec.
How do I run K2 Horizon 3.7B locally?
No Ollama library tag yet. Run the official GGUF with llama.cpp: the files declare the k2-horizon architecture, which mainline llama.cpp supports; an older Ollama or LM Studio build may not load it. The model card's own serving route is vLLM or SGLang. Running the published tag would send your prompts to a hosted GPU rather than your own machine.
What other sizes does K2 Horizon come in?
K2 Horizon 3.7B (4 GB), K2 Horizon 7B (6 GB), K2 Horizon MoVA 36B-A4B (23 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.