Kimi Delta Attention utilizes bra-ket notation for clarity in mathematical derivations, where column vectors represent states, row vectors represent duals, numbers represent scalars, and matrices represent linear transformations. In this notation, vectors are oriented to the right by default, while keys are oriented to the left within the linear attention structure.
blog.doubleword.ai
13 min
7/28/2026
Kimi Linear is a hybrid linear attention architecture that outperforms full attention in short-context, long-context, and reinforcement learning scenarios. It utilizes Kimi Delta Attention to achieve this efficiency.
arxiv.org
2 min
7/28/2026
Transformers utilize a query, key, and value (QKV) attention formulation that is crucial for AI tasks. The study investigates the individual contributions of these three projections and the effects of omitting any of them.
arxiv.org
2 min
6/4/2026
Transformer-based large language models struggle with long-context tasks due to poor scaling of their attention mechanism. Implementing a sleep-like consolidation mechanism allows models to convert recent context into persistent fast weights while clearing their key-value cache.
arxiv.org
2 min
5/26/2026
Kimi Delta Attention utilizes bra-ket notation for clarity in mathematical derivations, where column vectors represent states, row vectors represent duals, numbers represent scalars, and matrices represent linear transformations. In this notation, vectors are oriented to the right by default, while keys are oriented to the left within the linear attention structure.
blog.doubleword.ai
13 min
7/28/2026
Transformers utilize a query, key, and value (QKV) attention formulation that is crucial for AI tasks. The study investigates the individual contributions of these three projections and the effects of omitting any of them.
arxiv.org
2 min
6/4/2026
Kimi Linear is a hybrid linear attention architecture that outperforms full attention in short-context, long-context, and reinforcement learning scenarios. It utilizes Kimi Delta Attention to achieve this efficiency.
arxiv.org
2 min
7/28/2026
Transformer-based large language models struggle with long-context tasks due to poor scaling of their attention mechanism. Implementing a sleep-like consolidation mechanism allows models to convert recent context into persistent fast weights while clearing their key-value cache.
arxiv.org
2 min
5/26/2026
Kimi Delta Attention utilizes bra-ket notation for clarity in mathematical derivations, where column vectors represent states, row vectors represent duals, numbers represent scalars, and matrices represent linear transformations. In this notation, vectors are oriented to the right by default, while keys are oriented to the left within the linear attention structure.
blog.doubleword.ai
13 min
7/28/2026
Transformer-based large language models struggle with long-context tasks due to poor scaling of their attention mechanism. Implementing a sleep-like consolidation mechanism allows models to convert recent context into persistent fast weights while clearing their key-value cache.
arxiv.org
2 min
5/26/2026
Kimi Linear is a hybrid linear attention architecture that outperforms full attention in short-context, long-context, and reinforcement learning scenarios. It utilizes Kimi Delta Attention to achieve this efficiency.
arxiv.org
2 min
7/28/2026
Transformers utilize a query, key, and value (QKV) attention formulation that is crucial for AI tasks. The study investigates the individual contributions of these three projections and the effects of omitting any of them.
arxiv.org
2 min
6/4/2026
No more articles to load