A multimodal gap is revealed in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.
Abstract
Human language is highly polysemous. Many common words (e.g.,"bank"or"palm") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.
Theories of the mental lexicon must explain how people use context to interpret ambiguous words (e.g., "internal organ" vs. "musical organ") and why people vary on this core aspect of comprehension. Theoretical development has been hampered by the lack of reliable tests of disambiguation skill. We introduce a task in which participants hear three-sentence narratives with an ambiguous final word (e.g., "organ") and select the appropriate picture from two alternatives. Experiment 1 (N=197) confirmed the viability of this task using a repeated-measures design optimized for detecting group-level ambiguity effects: participants were slower and less accurate for ambiguous compared with closely matched control items in which the final word was replaced with a similar but unambiguous word, e.g., "piano"). We replicated this ambiguity disadvantage using a less closely matched control condition that was developed for Experiment 2 (N=242) which was optimised to measure variation in disambiguation skill. Experiment 2 confirmed robust variability in accuracy and response time, but measurement reliability ranged from r=.41 to .65, below the conventional r=.70 threshold, and correlations with other cognitive tasks were inconsistent. These results highlight the challenges of adapting complex, repeated-measures experimental paradigms to detect individual differences, and provide a framework for developing and assessing such measures.
L. M. Blott, A. Gowenlock, A. J. Parker et al.· Quarterly Journal of Experim...· 0 citations
Abstract Language models (LMs) are commonly criticized for not “understanding” language. However, many critiques focus on cognitive abilities that, in humans, are distinct from language processing. Here, we instead study a kind of understanding tightly linked to language: inferring “who did what to whom” (thematic roles) in a sentence. Does the central training objective of LMs—word prediction—result in sentence representations that capture thematic roles? In two experiments, we characterized sentence representations in four LMs that have been proposed as models of human language processing. The overall representational similarity of sentence pairs did not reflect whether they had the same agent/patient assignments or opposite agent/patient assignments. Furthermore, we found limited evidence that thematic role information was available in any subspace of hidden activations. However, some attention heads robustly captured thematic roles, independently of syntax. Therefore, LMs can extract thematic roles but this information influences their representations weakly.
Joseph M. Denning, Xiaohan Guo, Bryor Snefjella et al.· Annual Meeting of the Cognit...· 1 citation
Why what is really a matter of data analytics and statistical prediction is so readily assumed to be a display of real intelligence and even emergent cognition is explored by genealogically tracing the relationship between machines, organisms and language.
Chantelle Gray· Deleuze and Guattari Studies· 0 citations
While acoustic details such as voice onset time can be informative as to the identity of a phoneme (e.g., GOAT versus COAT), these details are not always reliable. Fortunately, disambiguating cues can often be found elsewhere, such as the semantic context. Indeed, acoustically ambiguous phonemes can be perceptually biased toward a semantically congruent percept (e.g., “The girl milked the ?OAT,” resulting in the percept GOAT). However, these findings come from well-controlled studies where participants repeatedly attend to specific contrasts produced in isolated sentences by a hyper-articulate speaker, which may limit generalizability to real-world listening tasks. Therefore, the current project aims to reconsider this research question in a more naturalistic listening scenario. A two-speaker referential communication task was used to elicit English discourse about visual scenes. Subsequently, statements containing critical minimal-pair referents were replaced with experimental trials containing acoustic × semantic manipulations. Participants in the current study will be told to listen to the conversation and show they’re paying attention by clicking referents in the on-screen image. Supposedly, their primary task will be to complete intermittent survey items about the speakers’ social communication. In reality, we will use eyetracking to infer their phonological and lexical processing during the embedded trials.
Mel Mallard, K. V. Van Engen· Journal of the Acoustical So...· 0 citations
This study investigates how the concept of morphological transcendence (
Libben, 2010
,
2014
) is manifest across semantic, visual, and phonetic
processing of English compounds. The theory of morphological transcendence proposes that constituents inside compounds take on
position-specific meanings that differ from their meanings as independent words. We first assess this claim using the CAOSS model
(
Marelli et al., 2017
), which represents a compound as the sum of two transformed
constituent semantic vectors that can be interpreted as transcended meanings. CAOSS provides a useful computational formalisation
of this idea and captures aspects of semantic transparency within its training data. However, it generalises poorly to unseen
compounds, highlighting the context sensitivity of compound interpretation and the limits of a single global mapping from
constituent meanings to compound meanings. We next examine subliminal reading using a Discriminative Lexicon Model (DLM;
Heitmeier et al., 2026
). Because reading is incremental, early comprehension depends on
the location of ocular fixation. We show that access to low frequency boundary trigrams allows the DLM to shift rapidly toward the
compound meaning, an effect we term ocular transcendence. Predictors from the DLM model account well for lexical decision
latencies, whereas predictors from the CAOSS model do not, suggesting that the two models capture different aspects of
comprehension. Finally, we analyse the pitch contours of three-constituent compounds. We demonstrate that these contours transcend
those of the constituents alone, providing evidence for an effect we term phonetic transcendence. Together, these findings support
a view of the lexicon as a dynamic, context-sensitive system, aligning with
Libben’s
(2022)
broader theoretical framework in which the mental lexicon is best understood as a flexicon.
M. Bell, R. Harald Baayen· The Mental Lexicon· 0 citations