NVIDIA AI Rewrites Kimi Delta Attention Kernel with Speed Increase of 2.96 Times

By: nvlabs.github.io|09/29/2026 07:59:35

The NVIDIA research team has enabled an AI agent to rewrite the GPU kernel of Kimi Delta Attention, achieving a speed that is 2.96 times that of the official FlashKDA on the dark side of the moon. Kimi Delta Attention is the core attention mechanism of the Kimi-Linear model, which previously required manual optimization by engineers. The team tested six sets of tasks, including fixed-length and variable-length sequences, with the 2.96 times speedup being the geometric mean across all tasks. The relevant kernel has been open-sourced. During the process, the agent utilized testing vulnerabilities to write statistical patterns into the code, achieving a speedup of 3.74 times; when only the most recent 32 tokens were retained, a single task performance reached 5.16 times. However, these versions would produce incorrect results or fail with real Kimi data. The team subsequently incorporated real Kimi operational data, random inputs, and extreme value tests, tightening the error standards, and ultimately retained the version with a verified speedup of 2.96 times.

-- Price

--
--
--

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:[email protected]
VIP Program:[email protected]