Skip to content
Preprint

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

Sep 2026 · 0 citations · 80 references
Computer Science

TL;DR

Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.

Abstract

State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.

View source

Similar papers

Preprint Sep 2026

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 1...

Kian Hosseinkhani, Qin-He Peng, George Shramko et al. · 1 citation
Preprint Aug 2026

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

FlashDrive is proposed, an algorithm-system co-design framework that targets all four stages of Vision-Language-Action inference simultaneously and moves end-to-end autonomous driving substantially closer to real-time deployment.

Ze-Kai Li, Yihao Liang, Hong-Fei Zhang et al. · 2 citations
#machine learning Preprint Sep 2026

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory, achieves state-of-the-art results on four memory benchmarks without loss on general-purpose control, and with the same backbone and training setup it outperforms retrieval, compression and recurrent-s...

C. Yin, Wang Xu, Jun-Peng Yang et al. · 1 citation
Preprint Sep 2026

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

D$^2$-VLA is presented, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA, which uses block-wise causal KV caching to encode observations incrementally and constructs separate historical KV read views for the VLM and action expert.

Zi-Jian Ye, Chen Wei, Wei Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.