Skip to content
Open access

The effect of syntactic and semantic information on word grounding through visual perception

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 31 references

Abstract

Word grounding refers to the ability of agents to associate linguistic terms with the perceptual concepts they represent, such as color, shape, and spatial features. This study quantitatively evaluates how syntactic and semantic information affects word-grounding performance. We evaluated five configurations within a generative multimodal Bayesian framework. The baseline grounds words using perceptual cues, while the additional configurations independently incorporate word-position indices, part-of-speech tags, FastText embeddings, and transformer-based contextual embeddings. The models were evaluated using a CLEVR-derived 3D scene dataset and synonym-substituted descriptions. Linguistic information improved grounding accuracy across the Color, Geometry, and Spatial Relation modalities. Contextual embeddings produced a 53.8% absolute improvement in object-grounding accuracy over the baseline. Only models incorporating semantic embeddings generalized effectively to synonym-substituted descriptions without retraining. These findings demonstrate the contribution of linguistic structure and semantic representations to perceptual alignment and provide empirical foundations for developing language-grounding systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.