Skip to content

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Sep 2026 · 1 citation · 68 references
Computer Science

TL;DR

SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory, achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-state methods by a wide margin.

Abstract

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, must decide what to keep from the past before knowing what a future decision will require. They were motivated by the assumption that minute-scale history is too large to process directly, which no longer holds for modern VLM backbones. We propose SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory. It keeps the sampled history intact in the timestamped video format the backbone was pretrained to process, routes the evidence it finds to a standard flow-matching action head through the hidden states of a generated sub-task, and prefills the history shared by consecutive decisions during action execution, keeping latency close to that of a single-frame VLA. SimpleMemVLA achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-state methods by a wide margin. History interventions show that the policy reads specific evidence from its past and follows edited histories without parameter updates, a visual form of in-context learning. On a physical dual-arm robot, it completes two tasks whose decisive evidence disappears before the robot acts. Code available at https://github.com/OpenBMB/SimpleMemVLA

View source

Similar papers

Preprint Sep 2026

SmoLSTM: A Compact Vision-Language-Action Model with Recurrent Memory that Persists

SmoLSTM couples a frozen 256M-parameter SmolVLM backbone to a matrix-memory LSTM control layer in which observation tokens and action queries are unified in a single causal stream that is never reset throughout the entire episode, capturing the episode in a recurrent state.

Jan-Gerrit Habekost, Parsa Mastouri Kashani, Connor Gäde et al. · 0 citations
#artificial intelligence Preprint Sep 2026

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but m...

Zi-Han Chen, Xuejian Rong, Xiao-Juan Wang et al. · 0 citations
Preprint Sep 2026

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.

Kemal Oksuz, Alexandru Buburuzan, Yu-Han Yao et al. · 0 citations
Preprint Aug 2026

UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

UniMem is presented, a framework that unifies high-level, multimodal memory and low-level control under one backbone that outperforms fixed-interval image sampling baselines in simulation and hierarchical baselines in hardware, while offering faster inference and a simple training pipeline for easy adoption.

Lars W. Osterberg, M. Wang, Mac Schwager · 1 citation
Preprint Aug 2026

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

VideoHarness-RSI is introduced, a controlled framework that recursively searches executable context constructors around a frozen VLM while keeping the answering model and interface fixed and establish executable context construction as a distinct optimization layer and provide an auditable baseline for studying harness...

Guo-Yang Xu, Hao Chen · 2 citations · ⚡1

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.