Skip to content
Review Open access

Vision‐Language‐Action Models for Embodied Artificial Intelligence: A Comprehensive Survey

Aug 2026 · Journal of Field Robotics · 0 citations · 31 references

Abstract

Embodied artificial intelligence (Embodied AI) aims to develop agents that can perceive, reason, and act through continuous interaction with physical world. To achieve this goal, agents must not only understand visual scenes and language instructions, but also transform multimodal information into executable actions. This requirement calls for models that can integrate perception, semantic reasoning, and action generation within a unified decision‐making framework. Recent advances in large language models (LLMs) have provided important technical foundations for this integration, leading to emergence of vision‐language‐action (VLA) models as a central paradigm for embodied systems. Although existing surveys have reviewed Embodied AI from perspectives such as robotic systems, simulation environments and task settings, few have systematically analyzed action‐generation mechanisms of VLA models. This limitation makes it difficult to distinguish VLA architectures. To address this gap, we provide a comprehensive review of VLA models for Embodied AI from an action‐generation perspective. We first formulate VLA models as embodied decision‐making systems that map visual observations and language instructions to action sequences through interaction with dynamic environments. Building on this formulation, we review historical evolution of VLA models. Furthermore, we propose an action‐generation‐centered taxonomy that categorizes VLA models into three paradigms: direct policy learning, generative action modeling, and reasoning‐guided modeling. For each paradigm, we analyze representative methods, action‐generation mechanisms, advantages and limitations, thereby providing a coherent framework for comparing diverse VLA architectures. Beyond model architectures, we summarize key resources and benchmarking settings and discuss recent applications. Finally, we identify major challenges and future directions. By organizing VLA models around the central problem of action generation, we aim to provide a structured technical map for future research toward more general, reliable, and physically grounded Embodied AI.

Read PDF