Preprint
Aug 2026
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
DART (Decoded Attention over Recurrent sTates), which retains the chunk state contributions produced by the Mamba-2 chunked scan as chunk state memories, decodes token-conditioned keys and values from these memories, and performs state-memory attention (SMA) over the resulting KV pairs.
Yixiao Qian, Song Chen, Pengkai Wang et al.
· 0 citations