General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize...
Hao-Jian Huang, Ze-Xi Li, Ju-Hao Guo et al.· 0 citations
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator tha...
Bin Lin, Bo Zhao, Bo-Yang Wang et al.· 0 citations
UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.
Venus Team, Zhuo-Hang Cai, Hao-Xin Chen et al.· 1 citation
MAGA is introduced that re-allocates training signal according to the structured action and suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions.
Hang Yan, Zhangxuan Gu, Bei-Tong Zhou et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.