Preprint
Aug 2026
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
DyLaR is proposed, which first grounds a question in a short block of perception latents, and then adaptively decides whether to append reasoning latents before answering, and which improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query.
Hao-Tian Xia, Zilin Xiao, Junbo Zou et al.
· 0 citations