Preprint
Aug 2026
Retrofitting Linear Attention into Diffusion Language Models
This work introduces block-hybrid attention, which retains exact softmax attention within the active denoising block while applying linear attention over previous blocks, and shows that pretrained dLLMs can be efficiently linearized for faster inference.
Jinha Kim, Younghun Roh, Jaeyeon Kim
· 0 citations