Task-optimized neural networks reveal distinct contributions of specialized and broader visual learning to neural representations of face familiarity
Abstract
How neural activity across the ventral visual hierarchy supports face recognition is an open question. A long-standing debate asks whether face processing, particularly in fusiform cortex, relies on face-specific computations or representations shared with broader visual recognition. Here we combine source-resolved magnetoencephalography (MEG) with task-optimized neural networks as controlled computational models of visual experience. Rather than manipulating long-term expertise in human observers, we systematically vary learning objective on the model side—what the networks are trained to recognize—while holding architecture and loss function constant within model comparisons. We then ask which learned representational geometries best align with neural responses, where and when. We measured millisecond-resolved brain–model alignment across V1, lateral occipital cortex (LOC) and fusiform cortex while participants viewed familiar, unfamiliar and scrambled faces. The same stimuli were presented to seven CNN architectures trained for face-identity recognition (FR), object-category recognition (OR) or object categorization including a face category (Dual), alongside untrained controls. Familiarity produced a stage-dependent dissociation: in LOC, familiar faces showed earlier brain–model alignment than unfamiliar faces in the M170 range, an effect most consistent in models trained for face-recognition, whereas broader objectives produced more variable, architecture-dependent peak alignment latencies. In fusiform cortex they showed stronger alignment around the M200 range. This fusiform advantage was not uniquely associated with face-recognition training: dual- and object-trained models showed greater fusiform correspondence than face-trained models. Together, these findings support stage-dependent specialization, with training for face-identity recognition constraining intermediate-stage timing while later fusiform representations remain compatible with representational structure acquired through broader visual computations. Significance Statement Recognizing a familiar face feels immediate, yet it remains unclear which stages of visual processing are specifically shaped by learning individual identities. We combine millisecond-resolved MEG with task-optimized neural networks used as controllable models of visual experience, manipulating learning objective on the model side—that is, what the networks are trained to recognize and tracking when and where the resulting representations align with human brain activity. Familiarity advances representational alignment in lateral occipital cortex but strengthens later alignment in fusiform cortex. The earlier LOC effect was most consistent across architectures after face-identity recognition training, whereas the later fusiform effect also emerged under broader visual learning objectives. These findings support a stage-dependent account of face recognition that moves beyond a simple face-specific versus broader-visual-processing dichotomy.