This work proposes StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes.
Abstract
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory, achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-s...
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information availabl...
Tej Deep Pala, Navonil Majumder, Bryce Goh et al.· 0 citations
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current obser...
Ya-Xin Zhao, Dian-Ye Huang, Chen-Wei Wang et al.· 0 citations
MEMOBench is presented, a benchmark for process level memory evaluation in robotic manipulation and provides a diagnostic evaluation suite and training supervision for memory grounded robotic policies.
Hai-Yang Sun, Haoxiao Wang, Jun-Ming Chen et al.· 0 citations
HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation, is proposed, enabling data-efficient compositional generalization.
Yun-Hao Zhao, Zhen-Yang Ni, Haoyang Chen et al.· 0 citations
D$^2$-VLA is presented, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA, which uses block-wise causal KV caching to encode observations incrementally and constructs separate historical KV read views for the VLM and action expert.
Zi-Jian Ye, Chen Wei, Wei Huang et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.