The integration of prosody and semantics in non-literal speech: A voxel-wise encoding model approach using large language models
Abstract
Irony and sarcasm are complex forms of non-literal language that hinge on a misalignment between surface meaning and speaker intent, requiring listeners to integrate contextual, semantic, and prosodic cues. While prior neuroimaging studies have implicated a broad network—including the temporal cortex, the inferior frontal gyrus, and the medial prefrontal cortex—in the comprehension of ironic and sarcastic speech, the precise neural mechanisms underlying the integration of semantic and prosodic information remain unclear. In the present study, we addressed this gap by employing voxel-wise encoding models to systematically identify brain regions specifically involved in combining prosodic and semantic cues during non-literal language comprehension. Participants listened to naturalistic auditory dialogues in which both discourse context and target utterance semantics and prosody were systematically manipulated. We derived custom text embeddings using transformer-based models to capture context-sensitive semantic representations of ironic statements, alongside acoustic features characterizing affective prosody. Ridge regression models were fitted to predict BOLD responses at the voxel level using semantic, prosodic, and combined features, and we identified integration as voxels in which each modality contributed predictive information beyond the other, using a permutation-based conjunction test. The regions integrating prosody and semantics depended on whether discourse context was modeled: integration was confined to the bilateral temporal speech cortex when statements were encoded in isolation, but additionally engaged the left inferior frontal gyrus pars orbitalis (IFGorb) when each statement was weighted by its relevance to the preceding context. These findings indicate that the left IFGorb integrates prosody with context-dependent meaning, engaging beyond the temporal speech cortex specifically when comprehension requires combining semantic, prosodic, and contextual cues—as in irony and sarcasm. Author summary In our daily conversations, people often say the opposite of what their words literally mean. When someone is being ironic or sarcastic, listeners rely not only on what is said but also on how it is said—the tone of voice—and on the broader context. We wanted to understand how the brain brings these pieces together to recover a speaker’s true meaning. Using artificial intelligence tools, we modeled the meaning of each sentence and gave more weight to the words most strongly linked to the earlier context. We combined these representations of meaning with measurements of vocal tone and tested how well they predicted brain activity while people listened to short dialogues. By comparing models that used only meaning, only tone, or both, we pinpointed regions that respond specifically to the combination of the two. When each sentence was modeled on its own, only speech regions in the temporal lobe combined tone and meaning. But once we let the context reshape a sentence’s meaning, a higher-level region in the left frontal lobe also came into play. This suggests that interpreting irony and sarcasm relies on a frontal region that blends tone with meaning after context has shaped it.