Kimi Linear, Explained: The Attention Architecture That Beat Full Attention
Moonshot AI's Kimi Linear is the first architecture to beat full attention under a controlled comparison — while cutting KV cache 75% and decoding 6.3x faster at 1M tokens. Here's how Kimi Delta Attention actually works, transcribed from the open-source kernel: the state-update equation, why per-channel gating was the unlock, where the 75% comes from arithmetically, and what the paper does not prove.
Every transformer you have ever run pays the same tax. To generate token number 100,001, the model re-reads all 100,000 tokens that came before it. It keeps them in a structure called the KV cache, and that cache grows linearly with your context — every token you add makes every future token more expensive to produce.
For years the fix was obvious and unusable. Linear attention replaces the growing cache with a fixed-size memory: instead of remembering every token, keep a compressed summary and update it as you go. Constant memory, constant cost per token, no matter how long the context. The catch was that these models were reliably worse. They forgot things. They failed at retrieval. Every few years someone proposed a new variant, the benchmarks came back soft, and the industry went back to paying the tax.
Kimi Linear, from Moonshot AI's Kimi Team, is the paper that broke the pattern. Under a controlled comparison — same data, same recipe, same parameter count — their hybrid architecture did not merely approach full attention. It outscored it, while cutting KV cache by 75% and decoding up to 6.3× faster at a million tokens of context.This is a walkthrough of how it works: the actual state-update equation, why one small change to a gating mechanism mattered so much, where the 75% comes from arithmetically, what the paper proves, what it does not, and what has happened in the ten months since.
TL;DR
Kimi Linear interleaves three Kimi Delta Attention (KDA) layers with one full Multi-head Latent Attention (MLA) layer, over and over. KDA is a fixed-size recurrent memory that stores no KV cache at all; MLA is ordinary exact attention that does. Three layers in four therefore cache nothing — which is exactly where the 75% KV cache reduction comes from. KDA's one real innovation is channel-wise forgetting: where Gated DeltaNet gives each attention head a single forget dial, KDA gives every feature dimension its own. In the paper's fair comparison at 1.4T training tokens, this scored 51.0 on MMLU-Pro (vs 47.2 for full MLA) and 84.3 on RULER at 128k (vs 81.3), with 6.3× faster time-per-output-token at 1M context. The released checkpoints are 48B total / 3B active MoE models with a 1M context window.
What this article covers
1. The tax every transformer pays
Standard softmax attention — the mechanism in the original 2017 transformer and in nearly every model you can download today — has a simple, brutal property: every token attends to every other token. For a sequence of length T, that is T² pairwise comparisons.
During training this quadratic cost is painful but parallelizable. During decoding, the cost shows up somewhere more annoying: the KV cache. To avoid recomputing the whole history for every new token, the model stores each token's key and value vectors in GPU memory. Generating the next token means reading that entire cache.
Two consequences follow, and they are the whole reason this paper exists:
- Memory grows without bound. At long context the KV cache stops being a footnote and starts competing with the model weights for VRAM. This is why a model that "supports 1M context" often cannot actually be run at 1M context on your hardware — see our guide to context windows for why the advertised number and the usable number differ.
- Time-per-token grows too. Decoding is memory-bandwidth bound. A cache that doubles in size takes roughly twice as long to stream. Your tokens-per-second decays as the conversation gets longer.
Linear attention attacks both at once by refusing to keep the cache at all.
2. Why linear attention kept losing
Strip attention down to its algebra and you find that softmax attention computes a weighted sum over all past values. Remove the softmax nonlinearity and you can re-associate the matrix products so that the model maintains a single state matrix S — think of it as an associative memory, a key-to-value lookup table squeezed into a fixed grid of numbers. Each new token writes its key-value pair into S; each query reads from it.
The math is elegant and the failure mode is obvious in hindsight. S has a fixed capacity. A 128-dimensional key space gives you a 128×128 matrix, and that is all the memory you get whether the context is one page or one library. Write enough key-value pairs into it and they interfere. The model develops what practitioners politely call "recall deficits" and everyone else calls forgetting the thing you asked about.
Early linear attention (2020-era) simply summed every token's contribution into S forever. Nothing was ever removed, so the state saturated into mush. The obvious patch was a forget gate: multiply the old state by a decay factor before each write, so old information fades. That is essentially what Mamba and its relatives do, and it works far better — but it is still a blunt instrument. Decay is uniform. Everything fades at the same rate, regardless of whether it was a critical variable name or a filler word.
3. The delta rule lineage
The more interesting line of work replaced "fade everything" with a targeted edit borrowed from classical associative-memory theory: the delta rule.
The intuition is that of a correction. Before writing key k into memory, first look up what the memory currently returns for k. If it already returns something, subtract that stale value out, then write the new one. Instead of blindly accumulating, the model performs a read-modify-write — it edits the specific slot associated with that key and leaves the rest of the memory alone.
4. What KDA actually changes: one dial per channel
Here is the entire innovation, stated plainly.
In Gated DeltaNet, the forget gate is a scalar per attention head. At each timestep the head decides "retain 95% of what I know" and applies that single number uniformly to its whole state matrix.
In Kimi Delta Attention, the forget gate is a vector with one entry per feature dimension. Each channel of the key space decides its own retention rate independently. One dimension can hold a value almost indefinitely while the dimension next to it flushes immediately.
You do not have to take the paper's word for this — it is visible in the shape of the tensors in the open-source KDA kernel in flash-linear-attention. The reference implementation documents its gate argument as:
g (torch.Tensor):
Per-dimension decay gates (log-space) of shape [B, T, HV, K].
That trailing K — the head dimension — is the whole paper. Gated DeltaNet's equivalent tensor stops at [B, T, HV]: one number per head per timestep. KDA carries a full K-vector.
This sounds like a minor engineering tweak. It is worth being precise about why it is not. A linear-attention model's entire competitive weakness is capacity: a fixed state must decide what to keep. Per-head gating forces one decision for the whole head — retain everything a bit, or forget everything a bit. Per-channel gating turns that single decision into K independent decisions, so the model can learn to dedicate specific dimensions to long-lived information (a variable's type, a document's topic) and let others churn on the local token stream.
The paper frames this as making better use of "limited finite-state RNN memory." That is exactly right, and it is why the change is not cosmetic: it targets the precise bottleneck that made every previous linear attention model lose.
5. The equation, line by line
Here is the KDA state update. This is not a paraphrase — it is transcribed from the reference implementation naive_recurrent_kda in flash-linear-attention:
S_t = (I − β_t k_t k_tᵀ) · Diag(α_t) · S_{t−1} + β_t k_t v_tᵀ
o_t = S_tᵀ q_t
Reading it as a sequence of operations on the memory S:
| Term | What it does |
|---|---|
Diag(α_t) · S_{t−1} | Forget, per channel. Scale each row of the state by its own decay factor. This is the KDA change; in Gated DeltaNet the multiplier is one scalar. |
(I − β_t k_t k_tᵀ) · … | Erase. A Householder-style projection that subtracts whatever the memory currently associates with key k_t, clearing the slot before reuse. |
+ β_t k_t v_tᵀ | Write. Store the new key-value association. β_t controls how forcefully. |
o_t = S_tᵀ q_t | Read. The output is just a lookup of the query against the current memory. |
And the corresponding lines of the reference implementation, which map one-to-one:
for i in range(0, T):
q_i, k_i, v_i, g_i, b_i = q[:, i], k[:, i], v[:, i], g[:, i], beta[:, i]
S = S * g_i[..., None].exp() # forget, per channel
S = S + einsum('b h k, b h v -> b h k v',
b_i[..., None] * k_i,
v_i - (k_i[..., None] * S).sum(-2)) # erase stale, then write
o[:, i] = einsum('b h k, b h k v -> b h v', q_i, S) # read
Two details worth noting. The gate is stored in log space and exponentiated (g_i.exp()), which keeps the decay strictly in (0, 1) and makes the cumulative product over a long sequence numerically stable — you sum logs instead of multiplying hundreds of thousands of fractions toward zero.
Second, the gate's parametrization is lifted almost verbatim from the Mamba-2 state-space lineage. The kernel computes it as:
g = −exp(A_log) · softplus(g_raw + dt_bias)
Those names — A_log, dt_bias — are the state-space discretization vocabulary. So KDA is fairly described as a hybrid of two research lines: it takes Mamba's data-dependent, per-channel gate parametrization and applies it to a delta-rule state update rather than a plain accumulate-and-decay one. Neither ingredient is novel alone. The combination is.
6. Why fine-grained gating was hard to make fast
If per-channel gating is such an obvious improvement, why did nobody ship it?
Because it breaks the trick that makes linear attention trainable at scale.
Nobody trains these models with the token-by-token loop shown above — a sequential for loop over a million timesteps would leave a GPU almost entirely idle. Instead, training uses a chunkwise formulation: split the sequence into chunks of, say, 64 tokens, compute everything within a chunk as dense matrix multiplications (which saturate the tensor cores), and only pass the state sequentially between chunks. You get near-parallel throughput with exactly recurrent semantics.
That reformulation depends on the per-step transition matrix having a shape you can algebraically collapse across a whole chunk. In Gated DeltaNet, the scalar gate makes this easy — a scalar commutes with everything, so it factors cleanly out of the chunk algebra.
Replace the scalar with Diag(α_t) and it no longer factors. The transition becomes a Diagonal-Plus-Low-Rank (DPLR) matrix: a diagonal part (the per-channel gate) plus a rank-one part (the delta-rule erase). General DPLR transitions have a known chunkwise algorithm, but it is expensive — it burns a large constant factor of extra matrix multiplies per chunk, enough to erase the efficiency advantage you were chasing in the first place.
The paper's second contribution is the fix: KDA is not a general DPLR system but a constrained special case, in which the low-rank component is tied to the same key vector used by the gate. In the DPLR notation, KDA fixes a_t = β_t k_t and b_t = k_t ⊙ α_t, rather than letting the two vectors vary freely. Because both terms are built from k_t, large parts of the chunkwise derivation collapse and cancel, and the authors were able to write a bespoke kernel that costs far less than the general formulation — while staying closer to the classical delta rule than general DPLR does.
This is the part of the paper that is genuinely hardware-aware research rather than architecture search. The open-sourced kernel shows the machinery involved: a wy_fast.py implementing the WY representation for compactly composing many Householder transforms, separate intra-chunk and inter-chunk passes, a token-parallel variant, and a fused_recurrent.py used for single-token decoding.
The part that is easy to miss
The architecture idea — finer gating — is one line of math. The reason it shipped is the kernel. A more expressive model that trains 3× slower is not an improvement, it is a worse model at fixed compute. Constraining the DPLR form so that a fast chunkwise algorithm exists is what converted the idea into something you could actually pretrain. Papers get cited for the equation; the equation ships because of the Triton.
7. The architecture: three KDA layers, one MLA layer, no RoPE
KDA alone is not the model. Kimi Linear is a hybrid, and the authors are direct about why: a fixed-size state, however well gated, cannot guarantee exact recall of an arbitrary token from a million-token context. Some layers still need to look at everything.
So the model interleaves, in a strict repeating pattern:
- Three KDA layers — linear, fixed-state, no KV cache.
- One MLA layer — full, exact attention. Multi-head Latent Attention is the DeepSeek-lineage mechanism that compresses keys and values into a low-rank latent vector before caching, so even the layers that do cache are cheaper than standard multi-head attention.
The paper reports this 3:1 ratio as the empirical sweet spot from their ablations. And one further choice is more interesting than it first appears: the MLA layers use NoPE — no positional encoding at all. No RoPE, nothing. All positional information is delegated to the KDA layers, whose decay dynamics inherently encode recency: something that entered the state 500 tokens ago has been multiplied by 500 decay factors, and that is a position signal.
This is an elegant division of labour. The linear layers handle "where and when," and the global layers are freed to do pure content-based retrieval unbiased by any hand-designed distance function — which is plausibly part of why the hybrid extrapolates well to very long contexts.
8. Where the 75% comes from — and what it does not cover
That last line deserves its own section, because the headline claim is routinely misread.
"75% less memory" means 75% less KV cache. It does not mean 75% less VRAM. These are different quantities, and for anyone sizing hardware the difference is the whole ballgame:- Model weights are constant. They do not care how long your context is. For the released 48B checkpoint at Q4_K_M, that is roughly 29 GB, and Kimi Linear does nothing to shrink it.
- KV cache scales with context length. It is the part that explodes at 128k and becomes ruinous at 1M — and it is the part the 3:1 ratio cuts by three quarters.
At a 4k context, the KV cache is a rounding error and Kimi Linear's memory advantage is close to nothing. At 1M tokens, the KV cache is the dominant term and the saving is transformative. The advertised win is a long-context win, and it is proportional to how long your context actually is. The same caveat applies to the speed number: the 6.3× figure is measured at 1M tokens. At short context, the paper's own Figure 1(a) places Kimi Linear at roughly the same speed as full attention — the claim there is that it matches on speed while winning on quality, not that it is 6× faster at every length.
If you want to see how the two quantities interact for any model and card, our VRAM calculator separates weights from cache explicitly, and the methodology page publishes the formulas.
9. The results
The paper's core experiment is a controlled comparison. Three architectures, identical training recipe, 1.4T tokens each: full MLA (the full-attention baseline), GDN-H (a hybrid built on Gated DeltaNet), and Kimi Linear. Same data, same budget — the comparison an architecture paper is supposed to run and frequently does not.
| Benchmark | MLA (full attention) | GDN-H | Kimi Linear |
|---|---|---|---|
| MMLU-Pro (4k context) | 47.2 | 47.9 | 51.0 |
| RULER (128k context) | 81.3 | 80.5 | 84.3 |
| Decoding acceleration at 128k | 1× | ~3× | 3.98× |
Two things stand out. First, Kimi Linear wins the short-context test (MMLU-Pro at 4k) by 3.8 points over full attention — and short context is where linear attention has no efficiency excuse and historically loses on quality. Second, GDN-H, the hybrid without per-channel gating, lands below full attention on RULER. The gap between 80.5 and 84.3 is the clearest evidence the paper offers that fine-grained gating is doing the work, since GDN-H is otherwise the same hybrid idea.
On throughput, the paper reports time-per-output-token against MLA at increasing decode lengths:
| Decode length | Speedup vs MLA (TPOT) |
|---|---|
| 4k–128k | small, growing |
| 256k | 4.8× |
| 512k | 5.7× |
| 1M | 6.3× (1.84 ms vs 11.48 ms) |
The released checkpoints, Kimi-Linear-48B-A3B-Base and -Instruct, were trained on 5.7T tokens — considerably more than the 1.4T used for the controlled architecture comparison. That distinction matters when reading third-party leaderboard scores for the public weights: those reflect a much larger training run, not the ablation above.
10. What the paper does not prove
The result is strong and the release is unusually complete — kernels, weights, and vLLM integration all open. Three honest caveats belong alongside the headline.
The comparison ceiling was 48B. The controlled experiments run at 48B total / 3B active parameters. Architecture conclusions have a long history of not surviving a two-order-of-magnitude scale-up, and nothing in the paper itself demonstrates that KDA's advantage holds at frontier scale. (Section 11 is about what happened when someone tried.) "Beats full attention on everything" is stronger than the evidence. What the paper shows is that Kimi Linear leads across the benchmark suite they evaluated, under their recipe, at their scale. That is a meaningful and well-controlled result. It is not a proof that no full-attention model, on any task, does better — and benchmark suites systematically under-sample the exact capability linear attention is theoretically weakest at: precise multi-hop retrieval over long context. There is a live counterexample. MiniMax shipped linear attention at frontier scale, encountered multi-hop reasoning deficits that their own benchmarks had not surfaced, and reverted to full attention in their M2 model. That is the strongest available argument that fixed-state memory can fail in ways standard evaluations miss, and it is not refuted by Kimi Linear's numbers — it is a caution about what those numbers can and cannot detect.The hybrid design is itself a concession to this. If linear attention were sufficient, the model would be 4:0, not 3:1. Keeping one exact-attention layer in four is an admission that a fixed-size state cannot be trusted alone for guaranteed recall.
11. What happened next: KDA at 2.8 trillion parameters
The paper went up in October 2025. The most useful evidence about whether it holds up arrived afterwards.
Kimi K3, announced in July 2026, is a 2.8-trillion-parameter open-weight model — and it runs Kimi Delta Attention in three of every four layers, keeping exact full attention only in the fourth. The same 3:1 hybrid, scaled roughly 58× past the paper's largest experiment, alongside a further architectural addition Moonshot calls attention residuals. Reported gains include a substantial improvement in overall scaling efficiency over K2.That is the answer to the "does it survive scale-up?" question, in the only form that really counts: a lab betting its flagship on it. Linear attention has moved from a research curiosity to production infrastructure at the largest open scale yet — and the 48B model documented in this paper is the reference implementation that got it there.
The caveat from Section 10 has not fully dissolved. The efficiency claims are arithmetic and have day-zero vLLM and SGLang numbers behind them. The quality claims at 2.8T are checkable from public weights but have not been independently verified with the kind of shortcut-free evaluation that would settle the multi-hop question. Worth watching rather than assuming.
If you are tracking the Kimi family more broadly, our writeup of Kimi K2.6 as an autonomous coding agent covers the agentic side of the same lineage.
12. Running Kimi Linear yourself
Both checkpoints are open weights.
| Models | moonshotai/Kimi-Linear-48B-A3B-Base · moonshotai/Kimi-Linear-48B-A3B-Instruct |
| Size | 48B total parameters, 3B active (Mixture-of-Experts) |
| Context | 1,048,576 tokens |
| Training | 5.7T tokens (released checkpoints) |
| Requirements | Python ≥ 3.10 · PyTorch ≥ 2.6 · fla-core ≥ 0.4.0 |
With vLLM, which has supported the architecture since release:
vllm serve moonshotai/Kimi-Linear-48B-A3B-Instruct \
--port 8000 \
--tensor-parallel-size 4 \
--max-model-len 1048576 \
--trust-remote-code
Or through Transformers directly:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "moonshotai/Kimi-Linear-48B-A3B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
model_name, torch_dtype="auto", device_map="auto", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
messages = [{"role": "user", "content": "Is 123 a prime?"}]
input_ids = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
print(tokenizer.batch_decode(model.generate(inputs=input_ids, max_new_tokens=500))[0])
A warning on llama.cpp and GGUF. KDA is a genuinely new operator, not a variation on attention that existing kernels handle. Support in llama.cpp has trailed the release: community GGUF conversions exist, but running them has required building from an open pull request rather than a released binary. If your local setup is Ollama or LM Studio, check current support before planning around it — this is the standard tax on being early to a new architecture, and it resolves on the usual timeline of weeks to months.
What it costs in VRAM
Using this site's standard sizing formula — weights only, before KV cache and runtime overhead:
| Quantization | Weights (48B total) | Streamed per token (3B active) |
|---|---|---|
| Q4_K_M | 28.98 GB | 1.81 GB |
| Q5_K_M | 34.02 GB | 2.13 GB |
| Q6_K | 39.36 GB | 2.46 GB |
| Q8_0 | 51.00 GB | 3.19 GB |
Note the 16× gap between the two columns, and note that it has nothing to do with KDA — it is the Mixture-of-Experts structure. All 48B parameters must be resident in memory, because any expert may be needed; only ~3B are read per decode step. That is why this model needs a 32 GB card (or two 24 GB cards, or a large unified-memory Mac) to hold at Q4, yet decodes at a speed characteristic of a 3B model. Confusing residency with per-token traffic is the single most common error in local-LLM sizing — see our quantization guide for the full picture, or check a specific card on the can-I-run tool.
13. Why this matters if you run models locally
The interesting thing about Kimi Linear is not that it is fast. It is that it stacks two independent efficiency mechanisms that attack different bottlenecks:
- Mixture-of-Experts decouples model quality from per-token compute. You pay VRAM for 48B and bandwidth for 3B.
- Hybrid linear attention decouples context length from per-token cost. You pay for a fixed state instead of a growing cache.
Neither helps with the other's problem — MoE does nothing for long context, and KDA does nothing for weight footprint. Together they describe the model shape that consumer and prosumer hardware can actually run: large enough to be smart, cheap enough per token to be fast, and flat enough in context to make a million-token window mean something other than an out-of-memory error.
That is the practical thesis worth taking away. Not "transformers are dead" — the model still contains exact attention, deliberately, in every fourth layer. The claim is narrower and more useful: you can replace most of the attention in a transformer with a well-designed recurrent memory and lose nothing, provided you keep enough exact attention to guarantee recall, and provided the gating is fine-grained enough to use the fixed state well. Kimi Linear is the first result that makes that case with a controlled experiment, an open kernel, and downloadable weights.
Sources and further reading
Primary- Kimi Linear: An Expressive, Efficient Attention Architecture — the paper (arXiv 2510.26692), Kimi Team, Moonshot AI
- MoonshotAI/Kimi-Linear — official repository, model table, and deployment instructions
- KDA kernel in flash-linear-attention — the open-source implementation;
naive.pyis the readable reference version quoted above - Kimi-Linear-48B-A3B-Instruct and -Base — the released checkpoints
- Gated DeltaNet — the architecture KDA extends
- Paper page on Hugging Face
- Designing Hardware-Aware Algorithms with Kimi Linear — DigitalOcean's technical walkthrough
- Linear Attention at Frontier Scale: Kimi K3's KDA Claim, Fact-Checked — the skeptical read, including the MiniMax M2 counterexample
- Kimi Linear recipe — LMCache deployment notes
- llama.cpp pull requests — current status of local GGUF support
- VRAM calculator — separate weights from KV cache for any model and card
- Methodology — the sizing and throughput formulas used above
- LLM context windows explained — why the advertised context and the usable context differ
- Quantization explained — what Q4_K_M actually costs you
- Kimi K2.6 for autonomous coding — the agentic side of the Kimi lineage