The results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.
Abstract
How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition, segmentation, optic-flow processing and predictive world modeling. Temporal integration improved object representations, but most video recognition models generalized poorly when appearance was disrupted while motion structure was preserved. Humans and macaque IT remained robust. Notably, predictive world models combined strong cross-appearance generalization with the closest correspondence to IT, outperforming other video-modeling approaches in neural fidelity. Yet no model reproduced the cortical transformation from early appearance-dominated responses toward later appearance-invariant motion coding. These results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.
An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vis...
The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principles underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypothese...
Nastaran Darjani, S. Robert, Maryam Vaziri-Pashkam et al.· bioRxiv· 0 citations
Human vision is profoundly non-uniform. Spatial resolution, contrast sensitivity, colour discrimination decrease, and the appearance of visual features become increasingly distorted with retinal eccentricity, yet visual experience appears remarkably rich and stable across the entire visual field. This apparent paradox...
Amelia Beatson, Anna Metzger, M. Toscani· Journal of Eye Movement Rese...· 0 citations
Color is a defining feature of human vision, yet its integration with spatial structure across stages of the visual system is still not fully understood. Classical accounts assumed that color provides little spatial information, being represented coarsely and separately from luminance. Here, we show that the spatial se...
Song-Lin Qiao, Karl R. Gegenfurtner, Ying-Fan Liu et al.· Science Advances· 0 citations
It is found that alignment with human fMRI, EEG, and macaque electrophysiology is already largely present at initialization, when networks classify at chance, and reaches a plateau within one to five epochs; thereafter it changes only modestly while classification accuracy continues to climb to 75%.
H. Scholte, Niklas Müller, Julio Smidi et al.· bioRxiv· 0 citations
Deep neural networks have achieved remarkable success in visual recognition; however, learning stable and discriminative visual representations under diverse imaging conditions remains a persistent challenge. Existing approaches often rely on increasing model depth or complex attention mechanisms, which improve perform...
Yu Wang, Y. Todo· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.