Skip to content
Open access

Teaching LLMs to Infer Real-World Consequences through Embodied Semantic Grounding

2026 · Computers, Materials & Continua · Vol 89, pp. 1-10 · 0 citations · 26 references

TL;DR

Embodied Semantic Grounding (ESG) is proposed, a framework that equips LLMs with consequence-aware text representations that extends affordance-grounding with consequence-level representations of post-event environmental functionality.

Abstract

: Large Language Models (LLMs) have recently advanced in real-world commonsense reasoning, including understanding everyday object behaviors and inferring their attributes from text. However, they remain limited in reasoning about the real-world consequences of events, such as how object failures, obstructions, or structural changes affect the surrounding environment-especially without visual or sensorimotor input. Existing works like PIQA and NEWTON evaluate narrow sub-skills, such as whether an object action makes sense and whether object properties can be inferred, providing valuable benchmarks for commonsense and physical reasoning but offering limited evaluation of how events alter environmental functionality and downstream conditions. To address this gap, we propose Embodied Semantic Grounding (ESG), a framework that equips LLMs with consequence-aware text representations. ESG learns a consequence-grounded space by aligning event descriptions with affordance maps-structured representations of how an environment can be used or traversed after an event-capturing how structural changes modify environmental functionality. A Flan-T5-XL model is trained with a contrastive alignment objective to encode event descriptions into this space, for coherent prediction of consequences such as collapses, blockages, and environmental changes. Rather than introducing a new language-model architecture, ESG extends affordance-grounding with consequence-level representations of post-event environmental functionality. We evaluate ESG on a unified benchmark comprising PIQA, NEWTON, LIBERO-derived affordance text, and 2400 synthetic scenario-based tasks. Results show that ESG improves performance over baseline language models across commonsense reasoning and consequence-prediction benchmarks. Under structured affordance-map supervision, ESG improves zero-shot accuracy on PIQA and NEWTON and demonstrates improved performance on synthetic consequence-prediction scenarios designed to evaluate post-event environmental reasoning.

Read PDF

Similar papers

Preprint Sep 2026

VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation

World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent...

Jianan Wang, Hao-Quan Zhai, Si-Yang Zhang et al. · 0 citations
#small language model Conference Open access Jul 2026

Reasoning Without Inference Cost: Latent Semantic Scaffolding for Robot VLA Policies

Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; adding causal reasoning improves manipulation, but current methods pay for it at inference time-generating reasoning tokens or rolling out predicted future states at every step, a cost that compounds over long hori...

A. Li, Zhuo Li, Zhe-Lin Yang et al. · 3 citations
Preprint Sep 2026

Generalist Open-World Temporal Perception

This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA), positioning generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physica...

C. Sminchisescu · 0 citations
Preprint Sep 2026

RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots...

Aernaer Akelijiang, Jian-Nan Li, Zhi-Neng Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models

Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated perspective. Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. Recent vision-language models (VLMs) appear capable of spat...

Mimo Shirasaka, Haochen Zhang, Yonatan Bisk · 0 citations
#artificial intelligence Preprint Sep 2026

Do World Models Learn Global Understanding?

AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we fra...

A. Detkov, Matt W. Thomson · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.