Skip to content
Preprint

Primate vision reveals a missing principle for robust dynamic AI

Aug 2026 · 1 citation
Computer Science Biology

TL;DR

The results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.

Abstract

How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition, segmentation, optic-flow processing and predictive world modeling. Temporal integration improved object representations, but most video recognition models generalized poorly when appearance was disrupted while motion structure was preserved. Humans and macaque IT remained robust. Notably, predictive world models combined strong cross-appearance generalization with the closest correspondence to IT, outperforming other video-modeling approaches in neural fidelity. Yet no model reproduced the cortical transformation from early appearance-dominated responses toward later appearance-invariant motion coding. These results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.

View source

Similar papers

Open access Aug 2026

Systematic image perturbations reveal persistent gaps between human and machine vision

An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vis...

Mugihiko Kato, Biyu J. He · 1 citation
Open access Sep 2026

Divergent specializations for motion-driven representations in higher lateral and dorsal visual areas

The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principles underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypothese...

Nastaran Darjani, S. Robert, Maryam Vaziri-Pashkam et al. · 0 citations
Review Open access Sep 2026

How Does Peripheral Vision Feel Rich? A Selective Review

Human vision is profoundly non-uniform. Spatial resolution, contrast sensitivity, colour discrimination decrease, and the appearance of visual features become increasingly distorted with retinal eccentricity, yet visual experience appears remarkably rich and stable across the entire visual field. This apparent paradox...

Amelia Beatson, Anna Metzger, M. Toscani · 0 citations
Open access Sep 2026

A unifying principle of chromatic coding across biological and artificial systems

Color is a defining feature of human vision, yet its integration with spatial structure across stages of the visual system is still not fully understood. Classical accounts assumed that color provides little spatial information, being represented coarsely and separately from luminance. Here, we show that the spatial se...

Song-Lin Qiao, Karl R. Gegenfurtner, Ying-Fan Liu et al. · 0 citations
Open access Aug 2026

Brain alignment in deep neural networks emerges early and independently of object classification

It is found that alignment with human fMRI, EEG, and macaque electrophysiology is already largely present at initialization, when networks classify at chance, and reaches a plateau within one to five epochs; thereafter it changes only modestly while classification accuracy continues to climb to 75%.

H. Scholte, Niklas Müller, Julio Smidi et al. · 0 citations
Conference Sep 2026

A bio-inspired ventral–dorsal fusion module for enhancing visual representations

Deep neural networks have achieved remarkable success in visual recognition; however, learning stable and discriminative visual representations under diverse imaging conditions remains a persistent challenge. Existing approaches often rely on increasing model depth or complex attention mechanisms, which improve perform...

Yu Wang, Y. Todo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.