Behavioral and Brain Responses to Language Reflect Different Levels of Linguistic Representation
Abstract
Human language processing can be studied through both behavior and brain activity, yet it remains unclear whether these two data types reflect sensitivity to the same information. One influential view holds that both behavioral and neural responses are largely determined by processing effort, often estimated by word surprisal together with the context-independent properties of word frequency and length. At the same time, neural responses have been shown to encode richer aspects of linguistic content, including meaning. Here, we use neural network language models to operationalize these alternatives and systematically compare, within the same analytic computational framework, the predictive power of low-dimensional effort-based predictors and high-dimensional embedding representations that encode contextualized linguistic content, including meaning. Across 8 behavioral datasets and 5 neural datasets (4 fMRI and 1 ERP), we find that processing effort captures substantial variance in both behavioral and neural measures of language processing, in line with much previous work. However, for brain responses—but not for behavioral measures—embedding representations carry substantial predictive power beyond the estimates of processing effort. These results therefore suggest that neural data provide access to rich, high-dimensional dynamics of language comprehension, whereas behavioral data reflect a bottlenecking of these dynamics into a small set of theoretically motivated properties of contextualized linguistic input. Significance Statement Two research communities study language comprehension as it unfolds in real time: psycholinguists use behavioral measures, such as eye movements during reading, and neuroscientists measure brain activity. The two are rarely studied together, but evidence from both must be integrated into a unified theory of language processing. Here we analyze both brain and behavioral responses within a single framework based on language models, comparing two long-standing accounts of what drives responses to language: processing effort versus meaning and other features not reducible to effort. We find that behavior is dominated by effort, whereas brain responses also reflect meaning. Developing a unified theory requires both kinds of data, but with a clear understanding of which levels of representation each measure reflects.