Skip to content

Author

Wenxuan Song

We have 12 of 46 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...

Jia-Yi Chen, Wen-Xuan Song, Jing-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...

Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al. · 0 citations
#machine learning Preprint Sep 2026

ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a fr...

Shuai-Tao Zhou, Kai-Sheng Pang, Wen-Xuan Song et al. · 1 citation
Preprint Aug 2026

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.

Lishan Yang, Wen-Xuan Song, Xi Wang et al. · 5 citations · ⚡1

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTS is introduced, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research and examines how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS d...

Shuxiang Zhang, Yiting Yin, Wen-Xuan Song et al. · 0 citations
Preprint Aug 2026

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

Xpolicylab is released as shared infrastructure for reproducible policy comparison and standardized deployment across simulation and physical platforms for reproducible policy comparison and standardized deployment across simulation and physical platforms.

XPolicyLab Community, Tian-Xing Chen, Yue Chen et al. · 6 citations · ⚡1
Preprint Aug 2026

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.

Zheng Yang, Wen-Jie Zhang, Xiang-Yu Chen et al. · 1 citation
Jul 2026

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiao-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency and across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods.

Xiaomin Guo, Piao-Piao Jin, Jason Li et al. · 16 citations · ⚡2
Jul 2026

MoWorld: A Flash World Model

MoWorld is the first real-time interactive World Model built on the Neural Processing Unit (NPU) and can achieves up to 50 FPS in such the devices, enabling practical and efficient deployment at scale.

Team Moxin, Deyi Ji, Tianrun Chen et al. · 1 citation
Preprint Aug 2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.

Hao-Dong Yan, Jun-Feng Li, Jun-Jie He et al. · 2 citations
Preprint Aug 2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.

Hao-Dong Yan, Jia-Guang Zhu, Ming-Ming Jia et al. · 6 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.