General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani...
Zi-Jie Diao, Yi-Tong Chen, Si-Cheng Xie et al.· 0 citations
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture rai...
Hao-Yu Zhao, Zi-Hao Zhao, Tian-Yuan Deng et al.· 0 citations
The proposed hierarchical long-horizon VLA architecture with an explicit language-memory module improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.
It is shown that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficult...
Zhikai Ding, Ziyi Ye· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.