Preprint
Aug 2026
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.
Marek Hradil, Danae Sánchez Villegas
· 0 citations