Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for plann...
Yu-Han Guo, Jin-Ming Liu, Liang Xu et al.· 0 citations
To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require mem...
Jian-Guo Huang, Jin-Ming Liu, Qi-Yao Wang et al.· 0 citations
FAN is introduced, which achieves the highest performance and demonstrates consistent robustness, providing insightful guidance for building stable action representations in achieving effective lifelong VLA adaptation.
Yi-Jun Hong, Jia-Run Zhu, Xiao-Quan Sun et al.· 0 citations
This work presents Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis, and proposes OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding.
Liang Xu, Cheng-Qun Yang, Zili Lin et al.· 0 citations
This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.