Sep 2026· IEEE Transactions on Pattern Analysis and Machine Intelligence· Vol PP, pp. 1-14· 0 citations
Medicine
TL;DR
EgoMotion is proposed, a two-stage framework for vision-language-guided egocentric motion generation that achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.
Abstract
Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely underexplored due to the inherent complexity of first-person perception. In this work, we investigate Egocentric Vision-Language (Ego-VL) motion generation. This task requires synthesizing 3D human motion conditioned jointly on first-person visual observations and natural language instructions. We identify a critical optimization challenge in effectively transferring vision-language understanding to motion generation. Directly optimizing vision-language semantic learning and kinematic motion synthesis in an end-to-end manner can lead to optimization interference, limiting the quality of generated motions. To address this challenge, we propose EgoMotion, a two-stage framework for vision-language-guided egocentric motion generation. In the first stage, a vision-language model (VLM) learns motion-aware semantic representations from multimodal inputs through autoregressive motion token prediction. In the second stage, these learned VLM representations serve as expressive conditioning signals for a diffusion-based motion generator. By performing iterative denoising within a continuous latent space, the generator synthesizes physically plausible and temporally coherent trajectories. Extensive evaluations demonstrate that EgoMotion achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.
CL4D is the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions, and 4DVLM, a 4D vision-language model that conditions language generation on dynamic geome...
K. Hewagamage, I. Senavirathne, S. Amarasinghe et al.· 0 citations
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...
Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al.· 1 citation
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...
Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-mod...
Guo-Cun Wang, Kenkun Liu, Guo-Rui Song et al.· 0 citations
Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared g...
This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.
Jiayu Ding, Zhuo-Dong Liu, Lei Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.