From Softmax to Kimi Delta Attention: the “obvious” path you missed
Kimi Delta Attention (KDA) turns attention into a recurrent “fast-weight” state update. By starting from softmax attention, removing the softmax bottleneck, then adding delta-rule correction and diagonal (channel-wise) forgetting, KDA becomes an almost inevitable refinement of linear attention.