Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
DyLaR is proposed, which first grounds a question in a short block of perception latents, and then adaptively decides whether to append reasoning latents before answering, and which improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query.