Do Latent Representations of Deep Visual Architectures Follow the Population Manifold Hypothesis of the Primate AIT Cortex?
Abstract
In computational neuroscience, the statistical nature of primate visual responses has long served as a benchmark for efficient coding. Specifically, previous works demonstrated that in the primate Anterior Inferotemporal (AIT) cortex, population sparseness $(S_{p})$ significantly exceeds single-neuron sparseness $\left(S_{l}\right)$, demonstrating that while individual neurons respond to relatively simple features, the total available feature space is vast. In this work, we establish a comparative experimental framework to bridge the gap between biological neural responses and the internal representations of Vision Transformers (ViTs) and ResNet-50. We analyze layer-wise kurtosis dynamics across ViT-B/16, ViT-L/16, ViT-B/32, and ResNet-50 using 3000 images of ImageNet-1k validation and Caltech-101 datasets. Our results reveal a “Semantic Snap” that is not a fluke and that it is a robust architectural phenomenon of visual models. Notably, while ResNet-50 inverts the biological signature $(S_{l}>S_{p})$, high-capacity transformer-based models like ViT-L/16 achieve a Lehky ratio, mirroring the distributed manifold coding of the AIT cortex. We further identify a significant magnitude gap in representational bandwidth between attention-based and convolutional architectures. These findings, validated by Pareto tail analysis, robust t-statistics, and Lehky ratio calculations, provide a computational link between transformer scaling laws and the Population Manifold Hypothesis in biological vision.