Skip to content
Preprint

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

Aug 2026 · 0 citations · 82 references
Computer Science

TL;DR

A multimodal gap is revealed in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.

Abstract

Human language is highly polysemous. Many common words (e.g.,"bank"or"palm") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.

View source

Similar papers

Aug 2026

EXPRESS: Measuring Individual Differences in Word-Meaning Disambiguation.

Theories of the mental lexicon must explain how people use context to interpret ambiguous words (e.g., "internal organ" vs. "musical organ") and why people vary on this core aspect of comprehension. Theoretical development has been hampered by the lack of reliable tests of disambiguation skill. We introduce a task in which participants hear three-sentence narratives with an ambiguous final word (e.g., "organ") and select the appropriate picture from two alternatives. Experiment 1 (N=197) confirmed the viability of this task using a repeated-measures design optimized for detecting group-level ambiguity effects: participants were slower and less accurate for ambiguous compared with closely matched control items in which the final word was replaced with a similar but unambiguous word, e.g., "piano"). We replicated this ambiguity disadvantage using a less closely matched control condition that was developed for Experiment 2 (N=242) which was optimised to measure variation in disambiguation skill. Experiment 2 confirmed robust variability in accuracy and response time, but measurement reliability ranged from r=.41 to .65, below the conventional r=.70 threshold, and correlations with other cognitive tasks were inconsistent. These results highlight the challenges of adapting complex, repeated-measures experimental paradigms to detect individual differences, and provide a framework for developing and assessing such measures.

L. M. Blott, A. Gowenlock, A. J. Parker et al. · 0 citations
Open access Apr 2025

Do Language Models Know Who Did What to Whom?

Abstract Language models (LMs) are commonly criticized for not “understanding” language. However, many critiques focus on cognitive abilities that, in humans, are distinct from language processing. Here, we instead study a kind of understanding tightly linked to language: inferring “who did what to whom” (thematic roles) in a sentence. Does the central training objective of LMs—word prediction—result in sentence representations that capture thematic roles? In two experiments, we characterized sentence representations in four LMs that have been proposed as models of human language processing. The overall representational similarity of sentence pairs did not reflect whether they had the same agent/patient assignments or opposite agent/patient assignments. Furthermore, we found limited evidence that thematic role information was available in any subspace of hidden activations. However, some attention heads robustly captured thematic roles, independently of syntax. Therefore, LMs can extract thematic roles but this information influences their representations weakly.

Joseph M. Denning, Xiaohan Guo, Bryor Snefjella et al. · 1 citation
Aug 2026

The Cybernetic Order-word: Tensors, Tensions and LLM Vector Spaces

Why what is really a matter of data analytics and statistical prediction is so readily assumed to be a display of real intelligence and even emergent cognition is explored by genealogically tracing the relationship between machines, organisms and language.

Chantelle Gray · 0 citations
Review Aug 2026

Developing a naturalistic approach for an experimental question: Perception of acoustically ambiguous phonemes within semantically disambiguating discourse

While acoustic details such as voice onset time can be informative as to the identity of a phoneme (e.g., GOAT versus COAT), these details are not always reliable. Fortunately, disambiguating cues can often be found elsewhere, such as the semantic context. Indeed, acoustically ambiguous phonemes can be perceptually biased toward a semantically congruent percept (e.g., “The girl milked the ?OAT,” resulting in the percept GOAT). However, these findings come from well-controlled studies where participants repeatedly attend to specific contrasts produced in isolated sentences by a hyper-articulate speaker, which may limit generalizability to real-world listening tasks. Therefore, the current project aims to reconsider this research question in a more naturalistic listening scenario. A two-speaker referential communication task was used to elicit English discourse about visual scenes. Subsequently, statements containing critical minimal-pair referents were replaced with experimental trials containing acoustic × semantic manipulations. Participants in the current study will be told to listen to the conversation and show they’re paying attention by clicking referents in the on-screen image. Supposedly, their primary task will be to complete intermittent survey items about the speakers’ social communication. In reality, we will use eyetracking to infer their phonological and lexical processing during the embedded trials.

Mel Mallard, K. V. Van Engen · 0 citations
Open access Aug 2026

A flexicon without words

This study investigates how the concept of morphological transcendence ( Libben, 2010 , 2014 ) is manifest across semantic, visual, and phonetic processing of English compounds. The theory of morphological transcendence proposes that constituents inside compounds take on position-specific meanings that differ from their meanings as independent words. We first assess this claim using the CAOSS model ( Marelli et al., 2017 ), which represents a compound as the sum of two transformed constituent semantic vectors that can be interpreted as transcended meanings. CAOSS provides a useful computational formalisation of this idea and captures aspects of semantic transparency within its training data. However, it generalises poorly to unseen compounds, highlighting the context sensitivity of compound interpretation and the limits of a single global mapping from constituent meanings to compound meanings. We next examine subliminal reading using a Discriminative Lexicon Model (DLM; Heitmeier et al., 2026 ). Because reading is incremental, early comprehension depends on the location of ocular fixation. We show that access to low frequency boundary trigrams allows the DLM to shift rapidly toward the compound meaning, an effect we term ocular transcendence. Predictors from the DLM model account well for lexical decision latencies, whereas predictors from the CAOSS model do not, suggesting that the two models capture different aspects of comprehension. Finally, we analyse the pitch contours of three-constituent compounds. We demonstrate that these contours transcend those of the constituents alone, providing evidence for an effect we term phonetic transcendence. Together, these findings support a view of the lexicon as a dynamic, context-sensitive system, aligning with Libben’s (2022) broader theoretical framework in which the mental lexicon is best understood as a flexicon.

M. Bell, R. Harald Baayen · 0 citations