Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations....