LFM2.5 2.6B — VRAM, Speed & Local Setup

Written by Jakub Rusinowski · Last updated September 19, 2026

Model libraryLFM2.5 → LFM2.5 2.6B

Liquid AI's on-device agentic model: 2.69B parameters over 30 layers, of which only 8 are attention — the other 22 are double-gated short convolutions that cache nothing. That is why its KV cache stays small at a 128K context where a conventional 2.7B model's would not. Trained on 34T tokens, post-trained for tool calling.

LFM2.5 2.6B needs about 2 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Specifications

Parameters2.69 Billion
Context window131,072
ArchitectureHybrid (22 conv + 8 GQA)
ProviderLiquid AI
LicenceLFM Open License v1.0
Specified atQ4_K_M
System RAM8 GB
Record updated2026-09-19

Corroborated — Two or more independent sources agree on these figures, but the model card itself was not retrieved. Treat the numbers as good rather than confirmed.

Licence

LFM Open License v1.0commercial use permitted. Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.

VRAM and Speed by Quantization

Modelled on a reference NVIDIA RTX 4090 (24 GB). Weights plus framework overhead only — this model publishes no architecture we can read, so no KV cache is included. A real session needs more; the figure is a floor, not a target. Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.

QuantBits/weightWeightsVRAM neededEst. speedFit on 24 GB
Q2_K2.630.9 GB1.7 GB~253 tok/s (est.)Fits comfortably
Q3_K_M3.411.1 GB1.9 GB~233 tok/s (est.)Fits comfortably
Q4_K_M4.831.6 GB2.4 GB~203 tok/s (est.)Fits comfortably
Q5_K_M5.671.9 GB2.7 GB~189 tok/s (est.)Fits comfortably
Q6_K6.562.2 GB3 GB~176 tok/s (est.)Fits comfortably
Q8_08.502.9 GB3.7 GB~152 tok/s (est.)Fits comfortably
F1616.005.4 GB6.2 GB~101 tok/s (est.)Fits comfortably

Want to set your own context length and KV-cache quantization? Use the interactive VRAM calculator.

Buy This HardwareIntel Arc B570 10GB — 10 GB VRAM · 150 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Recommended GPU

The cheapest catalogued GPU that runs LFM2.5 2.6B is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.
Check price on Amazon

How to Run LFM2.5 2.6B

Install Ollama, then run:

ollama run lfm2-5

Weights on Hugging Face: LiquidAI/LFM2.5-2.6B.

Best for: on device, agents, tool calling, edge.

Can I Run LFM2.5 2.6B on My GPU?

Other LFM2.5 Sizes

LFM2.5 2.6B — Frequently Asked Questions

How much VRAM does LFM2.5 2.6B need?
About 2 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does LFM2.5 2.6B run on an RTX 4090 (24 GB)?
Yes. LFM2.5 2.6B needs about 2 GB at Q4_K_M, inside a 24 GB card, at an estimated 203 tokens/sec.
How do I run LFM2.5 2.6B locally?
Install Ollama and run `ollama run lfm2-5`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does LFM2.5 come in?
LFM2.5 2.6B (2 GB), LFM2.5-8B-A1B (6 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.

← All LFM2.5 models | VRAM calculator | Check your own hardware