Preprint
Aug 2026
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
It is found that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.
David Bamman, Kent K. Chang, Allison Cooper et al.
· 2 citations