Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the crit...
Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...
Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al.· 0 citations
NavGen is introduced, a text-to-video data generation pipeline that produces diverse vision-language navigation episodes across indoor and outdoor scenes and a style-diversification method that scales up long-tail data that is difficult and costly to collect.
Xi-Jie Huang, Yong-Yang Wan, Cheng-Bin Dong et al.· 0 citations
A framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface is proposed, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training.
Yuan Zhou, Ruitong Lin, Shen Wang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.