Preprint
Aug 2026
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
Variance Recovery Policy Optimization (VRPO) is introduced, which retains and progressively expands groups to recover informative signals from prompts that are difficult yet solvable and retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable.
Jingqi Tian, Haoji Zhang, Lin Chen et al.
· 0 citations