Kimi Linear: An Expressive, Efficient Attention Architecture
摘要
论文提出 Kimi Linear,一种混合线性注意力架构,核心是 KDA 模块,扩展了 Gated DeltaNet,通过更细粒度门控和 DPLR 矩阵的专用变体提高硬件效率。在 3B 激活参数、48B 总参数的模型上,与全 MLA 相比,在相同训练配方下全面超越,KV 缓存减少 75%,1M 上下文解码吞吐提升 6 倍,并开源了 KDA 内核、vLLM 实现和模型权重。
荐读理由
这篇论文提供了可复现的开源实现和明确的性能数据,独立开发者可以直接评估KDA内核或vLLM集成,用于长上下文推理场景,减少KV缓存成本;同时它挑战了线性注意力不如全注意力的常见认知,值得深入验证。
原文
Computer Science > Computation and Language
[Submitted on 30 Oct 2025 (v1), last revised 1 Nov 2025 (this version, v2)]
Title:Kimi Linear: An Expressive, Efficient Attention Architecture
Authors:Kimi Team: Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, Chengyin Liu, Xin Men, Songlin Yang, Zhiyuan Li, Wentao Li, Enzhe Lu, Weizhou Liu, Yanru Chen, Weixin Xu, Longhui Yu, Yejie Wang, Yu Fan, Longguang Zhong, Enming Yuan, Dehao Zhang, Yizhi Zhang, T.Y. Liu, Haiming Wang, Shengjun Fang, Weiran He, Shaowei Liu, Yiwei Li, Jianlin Su, Jiezhong Qiu, Bo Pang, Junjie Yan, Zhejun Jiang, Weixiao Huang, Bohong Yin, Jiacheng You, Chu Wei, Zhengtao Wang, Chao Hong, Yutian Chen, Guanduo Chen, Yucheng Wang, Huabin Zheng, Feng Wang, Yibo Liu, Mengnan Dong, Zheng Zhang, Siyuan Pan, Wenhao Wu, Yuhao Wu, Longyu Guan, Jiawen Tao, Guohong Fu, Xinran Xu, Yuzhi Wang, Guokun Lai, Yuxin Wu, Xinyu Zhou, Zhilin Yang, Yulun Du
View a PDF of the paper titled Kimi Linear: An Expressive, Efficient Attention Architecture, by Kimi Team: Yu Zhang and 58 other authors
Abstract:We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.
https://doi.org/10.48550/arXiv.2510.26692
arXiv-issued DOI via DataCite
| Comments: | |
|---|---|
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2510.26692 [cs.CL] |
| (or arXiv:2510.26692v2 [cs.CL] for this version) | |
Submission history
From: Yulun Du [view email] [v1] Thu, 30 Oct 2025 16:59:43 UTC (645 KB) [v2] Sat, 1 Nov 2025 12:05:18 UTC (691 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Kimi Linear: An Expressive, Efficient Attention Architecture, by Kimi Team: Yu Zhang and 58 other authors
Current browse context:
cs.CL
Change to browse by:
References & Citations
Loading...
BibTeX formatted citation
Data provided by:
Bookmark
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Connected Papers (What is Connected Papers?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
alphaXiv (What is alphaXiv?)
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Hugging Face (What is Huggingface?)
ScienceCast (What is ScienceCast?)
Demos
Replicate (What is Replicate?)
Hugging Face Spaces (What is Spaces?)
TXYZ.AI (What is TXYZ.AI?)
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
这条对你有帮助吗?