Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.
Abstract
State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.
Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 1...
Kian Hosseinkhani, Qin-He Peng, George Shramko et al.· 1 citation
FlashDrive is proposed, an algorithm-system co-design framework that targets all four stages of Vision-Language-Action inference simultaneously and moves end-to-end autonomous driving substantially closer to real-time deployment.
Ze-Kai Li, Yihao Liang, Hong-Fei Zhang et al.· 2 citations
SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory, achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-s...
These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules, rather than universal layer-sensitivity rules for vision-language-action models.
Jiu-Yi Xu, Qing Jin, Mei-Da Chen et al.· 0 citations
D$^2$-VLA is presented, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA, which uses block-wise causal KV caching to encode observations incrementally and constructs separate historical KV read views for the VLM and action expert.
Zi-Jian Ye, Chen Wei, Wei Huang et al.· 0 citations
This work proposes StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes.
Wen-Zhuo Li, Q. Shi, Yi Zhou· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.