Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions typically extract independent memory entries from fixed-length video clips and thus cannot capture high-level semantics that need to be summa...
Yi-Zhou Tian, Zi-Zhe Chen, Shi-Yuan Deng et al.· 0 citations
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize...
Hao-Jian Huang, Ze-Xi Li, Ju-Hao Guo et al.· 0 citations
Robo-Harness K1 is introduced, a robot-use agent (RUA) framework that exposes perception as tools that makes 3D geometry accessible without changing the VLM architecture or training a depth encoder, and suggests that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies t...
Ze-Xi Li, Ye-Hang Zhang, Wen-Qian Li et al.· 0 citations
World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
Ye-Hang Zhang, Hao-Jian Huang, Yi-Fan Chang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.