Zeva is presented, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen, and achieves the best performance among the compared frontier VLAs and WAMs and enables self-evolution during deployment without gradient updates.
Abstract
Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.
Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...
Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al.· 0 citations
GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in co...
This work proposes an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions, and introduces a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to impr...
Shi-Lin Ma, Chu-Bin Zhang, Xu-Long Bai et al.· 0 citations
Motus2 is presented, a self-evolving general world model for dexterous manipulation that combines egocentric data scaling and closed-loop general world model scaling to provide a general path toward self-evolving dexterous manipulation.
Hong-Zhe Bi, Zikun Zhou, Yihao Tang et al.· 4 citations
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder...
Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al.· 0 citations
In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into en...
Chu-Han Zhang, E. Shahabi, K. Khomenko et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.