A vision-language-action (VLA) policy with a flow-matching action expert generates each action chunk (a short command sequence) by integrating a learned velocity field; once its weights are fixed, the success or failure of an earlier rollout cannot change the chunk generated now. Concurrent test-time methods give a fro...
Jia-Xuan Zhang, Rui-Zhe Liu, Yu Zhang et al.· 1 citation
Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile...
Zi-Hao Wang, Shu-Tong Liu, Si-Qi Zheng et al.· 0 citations
SlackDrive is proposed, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator.
Xiao-Huan Pei, Heng-Guang Zhou, Yuan-Hao Ban et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.