Skip to content
Preprint

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Aug 2026 · 3 citations · 43 references
Computer Science

TL;DR

Zeva is presented, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen, and achieves the best performance among the compared frontier VLAs and WAMs and enables self-evolution during deployment without gradient updates.

Abstract

Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.

View source

Similar papers

Preprint Sep 2026

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...

Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al. · 0 citations
Preprint Aug 2026

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in co...

Cheng-Hao Gu, Hanyang Yu, Jingbo Zhang et al. · 2 citations · ⚡2
Preprint Sep 2026

From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments

This work proposes an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions, and introduces a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to impr...

Shi-Lin Ma, Chu-Bin Zhang, Xu-Long Bai et al. · 0 citations
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder...

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations
Preprint Aug 2026

Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into en...

Chu-Han Zhang, E. Shahabi, K. Khomenko et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.