Kolibri — local AI model by Aleph Alpha
Written by Jakub Rusinowski · Last updated
Kolibri 1 is Aleph Alpha's mixture-of-experts reasoning model for German and English, released in October 2026 with Apache-2.0 weights. It was trained from scratch, supports adjustable reasoning effort and tool calling, and uses a mostly sliding-window attention design to keep long-context cache costs low. The memory trade-off is real: every expert must be resident, so it is a server and high-memory-workstation model, not a laptop model.
Variants
The smallest Kolibri variant needs about 48 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| Kolibri 1 → 78B (3.5B active) | ~48.3 GB | 262,144 | cloud-hosted tag |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run Kolibri locally
Install Ollama, then pull the tag.
No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory.
Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: Kolibri 1Recommended GPU
The cheapest catalogued GPU that runs Kolibri locally (min 48 GB VRAM) is the NVIDIA DGX Spark (128 GB).
Kolibri — frequently asked questions
How much VRAM does Kolibri need?
Kolibri needs about 48 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Kolibri 1 (48 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Kolibri on an RTX 4090 (24 GB)?
Kolibri's smallest variant needs about 48 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.
What quantization should I use for Kolibri?
Q4_K_M is the best balance of quality and VRAM for Kolibri in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Kolibri with Ollama?
Kolibri has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.