An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vision.
Abstract
Deep neural networks (DNNs) are promising computational models for understanding visual object recognition. Yet, whether DNNs use similar visual cues for object recognition as humans do remains unknown. We created an image set that systematically untangles global shape, internal parts, and texture information, and compared human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives. No DNNs replicated humans’ cue-reliance profile, including those with recurrence or specialized training. Fine-tuned text-image contrastive-trained models, regardless of architecture, were most human-like overall, but lost their human-alignment when the global shape was disrupted. Strikingly, all DNNs substantially underperformed humans when the global shape cue alone was critical to object recognition. Furthermore, alignment with ventral stream neural recordings in an existing database did not predict alignment to human behavior, and model performance does not always predict its human-alignment. Together, these findings reveal systematic and persistent differences between human and machine vision.
It is found that alignment with human fMRI, EEG, and macaque electrophysiology is already largely present at initialization, when networks classify at chance, and reaches a plateau within one to five epochs; thereafter it changes only modestly while classification accuracy continues to climb to 75%.
H. Scholte, Niklas Müller, Julio Smidi et al.· bioRxiv· 0 citations
The results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.
Matteo Dunnhofer, Christian Micheloni, Kohitij Kar· 1 citation
How neural activity across the ventral visual hierarchy supports face recognition is an open question. A long-standing debate asks whether face processing, particularly in fusiform cortex, relies on face-specific computations or representations shared with broader visual recognition. Here we combine source-resolved mag...
Hamza Abdelhedi, Shahab Bakhtiari, Karim Jerbi· bioRxiv· 0 citations
This work proposes BrainTrain, a framework to create more robust DNNs through human behavior alignment and shows its utility in the context of object recognition and proposes Similarity Driven Label Smoothing (SDLS), a regularization method that scales BrainTrain to applications where it is difficult or expensive to co...
Bharath Anand, Sarada Krithivasan· Frontiers in Artificial Inte...· 0 citations
Abstract Motivation The inductive bias of a deep learning model influences the features it extracts from biological images, making model selection a critical scientific decision. We systematically compare representations learned from scratch without pre-training by convolutional neural networks (CNNs), vision transform...
Jacob I. Evarts, Jason Y. Cain, Po-Hao Chiu et al.· Bioinformatics Advances· 0 citations
This work introduces axis-aligned feature accentuation, which converts each model’s fitted encoding axis into graded stimulus perturbations that are predicted to parametrically control neural firing within and beyond the natural-image range.
Jacob S. Prince, Binxu Wang, Thomas Fel et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.