Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privac...
Jun-Qiao Fan, Yuxuan Hu, Bofan Lyu et al.· 0 citations
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting pla...
Chuan-Liang Xie, Bo-Yu Ma, Gen Li et al.· 0 citations
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu...
Yun-Bei Zhang, Zi-Jian Jin, Yuan-Zhe Liu et al.· 0 citations
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information availabl...
Tej Deep Pala, Navonil Majumder, Bryce Goh et al.· 0 citations
Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on...
Shilin Shan, Chu-Hao Zhou, Rui-Ze Wang et al.· 0 citations
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalab...
Mengfei Zhao, Di-Hong Huang, Yikai Tang et al.· 0 citations
Real-world experiments demonstrate that a single $\omega$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.
It is shown that robots can generalize effectively under dynamics uncertainties even with limited training data by leveraging a feedback mechanism, namely PhyFilter, that corrects learning outputs with physics-filtered learning residuals, and shows that physics-filtered feedback can serve as a powerful alternative to m...
Jindou Jia, Shixu Han, Meng Wang et al.· npj Robotics· 1 citation
This work proposes UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation that unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-a...
Xin-Yu Zhou, Zikun Cai, Kuangji Zuo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.