Skip to content

Author

Jucheng Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Vision‐Language‐Action Models for Embodied Artificial Intelligence: A Comprehensive Survey

Embodied artificial intelligence (Embodied AI) aims to develop agents that can perceive, reason, and act through continuous interaction with physical world. To achieve this goal, agents must not only understand visual scenes and language instructions, but also transform multimodal information into executable actions. This requirement calls for models that can integrate perception, semantic reasoning, and action generation within a unified decision‐making framework. Recent advances in large language models (LLMs) have provided important technical foundations for this integration, leading to emergence of vision‐language‐action (VLA) models as a central paradigm for embodied systems. Although existing surveys have reviewed Embodied AI from perspectives such as robotic systems, simulation environments and task settings, few have systematically analyzed action‐generation mechanisms of VLA models. This limitation makes it difficult to distinguish VLA architectures. To address this gap, we provide a comprehensive review of VLA models for Embodied AI from an action‐generation perspective. We first formulate VLA models as embodied decision‐making systems that map visual observations and language instructions to action sequences through interaction with dynamic environments. Building on this formulation, we review historical evolution of VLA models. Furthermore, we propose an action‐generation‐centered taxonomy that categorizes VLA models into three paradigms: direct policy learning, generative action modeling, and reasoning‐guided modeling. For each paradigm, we analyze representative methods, action‐generation mechanisms, advantages and limitations, thereby providing a coherent framework for comparing diverse VLA architectures. Beyond model architectures, we summarize key resources and benchmarking settings and discuss recent applications. Finally, we identify major challenges and future directions. By organizing VLA models around the central problem of action generation, we aim to provide a structured technical map for future research toward more general, reliable, and physically grounded Embodied AI.

Ning Xiong, Mingle Xu, Wei Chen et al. · 0 citations