Preprint
Jul 2026
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
It is proved that the computationally cheaper split space-time attention is equivalent to full space-time attention and is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
N. Tran, Fanghui Xue, Shuai Zhang et al.
· 0 citations