Preprint
Aug 2026
LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding
LongVU-TTT is introduced, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates between the vision encoder and the LLM, and is stronger than attention- and fixed-state recurrent resamplers across three benchmarks.
Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase et al.
· 0 citations