Skip to content
Preprint

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

Aug 2026 · 2 citations · 33 references
Computer Science

TL;DR

FlashDrive is proposed, an algorithm-system co-design framework that targets all four stages of Vision-Language-Action inference simultaneously and moves end-to-end autonomous driving substantially closer to real-time deployment.

Abstract

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.

View source

Similar papers

Preprint Aug 2026

FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

FlashVLA is introduced, a streaming action decoding framework that addresses both challenges in a unified formulation and can achieve $\geq$30\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.

Ze-Kai Li, Jiarui Tang, Zhijian Liu · 4 citations
Preprint Sep 2026

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.

Kemal Oksuz, Alexandru Buburuzan, Yu-Han Yao et al. · 0 citations
Preprint Aug 2026

Inference-Time Attention Steering for Vision-Language-Action Driving Models

A bounded additive pre-softmax attention bias on the visual tokens of detector localized traffic actors on Alpamayo-R1's Qwen3-VL backbone is studied, suggesting the bias governs where the model looks rather than encoding a target behavior.

D. Prasad, Lars Ullrich, Knut Graichen · 1 citation
Preprint Aug 2026

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI is proposed, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters and seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference.

Zhe Liu, Jinghua Hou, Yuxiang Lu et al. · 1 citation · ⚡1
Preprint Sep 2026

Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but...

Yi-Heng Ji, Xing-Ru Zhou, Luis Sentis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.