Skip to content

StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

This work proposes StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes.

Abstract

Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.

View source

Similar papers

#machine learning Preprint Sep 2026

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory, achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-s...

C. Yin, Wang Xu, Jun-Peng Yang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information availabl...

Tej Deep Pala, Navonil Majumder, Bryce Goh et al. · 0 citations
Preprint Sep 2026

Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current obser...

Ya-Xin Zhao, Dian-Ye Huang, Chen-Wei Wang et al. · 0 citations
Preprint Aug 2026

Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation, is proposed, enabling data-efficient compositional generalization.

Yun-Hao Zhao, Zhen-Yang Ni, Haoyang Chen et al. · 0 citations
Preprint Sep 2026

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

D$^2$-VLA is presented, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA, which uses block-wise causal KV caching to encode observations incrementally and constructs separate historical KV read views for the VLM and action expert.

Zi-Jian Ye, Chen Wei, Wei Huang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.