DriveCache is proposed, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget, which improves the overall fidelity-efficiency trade-off over evaluated cache methods.
Jianchun Yang, Jian Liang, Xian-Da Guo et al.· 0 citations
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the same visual observation may imply different actions for different agents, making prediction ambiguous when relying solely on vision. Existing studies mainly rely on reinforcement learning, which requires large...
Jialu Zhang, Yong Du, Xianda Guo et al.· arXiv.org· 0 citations
OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision, and is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision.
Hong-Yuan Zhang, Xian-Da Guo, Yan-Lun Peng et al.· Information Fusion· 0 citations
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied navigator requires, producing recurring failures such as dead-end stalls, in-room loops, and circuitous approaches to detected targets. Prompt...
Rui Sang, Yiqun Duan, Pinhan Fu et al.· arXiv.org· 0 citations
DreamWAM is introduced, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics, showing that robust world-action learning depends not only on predicting the future, but on representing it in a fo...
Shanglin Yuan, Weiheng Zhao, Xin Shi et al.· 4 citations
TrustVLA is introduced, a mechanism-guided inference-time defense that adapts the Dirichlet evidence framework from trusted classification to monitor per-token, per-layer epistemic uncertainty in VLA policies, providing a retraining-free, mechanism-guided defense for visual-triggered VLA backdoors.
Pin-Han Fu, Xian-Da Guo, Xue-Tao Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.