Kolibri — local AI model by Aleph Alpha

Written by Jakub Rusinowski · Last updated

Kolibri 1 is Aleph Alpha's mixture-of-experts reasoning model for German and English, released in October 2026 with Apache-2.0 weights. It was trained from scratch, supports adjustable reasoning effort and tool calling, and uses a mostly sliding-window attention design to keep long-context cache costs low. The memory trade-off is real: every expert must be resident, so it is a server and high-memory-workstation model, not a laptop model.

Variants

The smallest Kolibri variant needs about 48 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.

ModelVRAM
Kolibri 1 →
78B (3.5B active)
~48.3 GB

Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.

How to run Kolibri locally

Install Ollama, then pull the tag.

No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory.

Pick a size above for its own VRAM figure, speed estimate and install command.

Licence

Apache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Applies to: Kolibri 1

Recommended GPU

The cheapest catalogued GPU that runs Kolibri locally (min 48 GB VRAM) is the NVIDIA DGX Spark (128 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
NVIDIA DGX Spark (128GB)
128 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.

Kolibri — frequently asked questions

How much VRAM does Kolibri need?

Kolibri needs about 48 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Kolibri 1 (48 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run Kolibri on an RTX 4090 (24 GB)?

Kolibri's smallest variant needs about 48 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.

What quantization should I use for Kolibri?

Q4_K_M is the best balance of quality and VRAM for Kolibri in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Kolibri with Ollama?

Kolibri has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.