Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting pla...
Chuan-Liang Xie, Bo-Yu Ma, Gen Li et al.· 0 citations
Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on...
Shilin Shan, Chu-Hao Zhou, Rui-Ze Wang et al.· 0 citations
This work proposes UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation that unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-a...
Xin-Yu Zhou, Zikun Cai, Kuangji Zuo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.