FlashDrive is proposed, an algorithm-system co-design framework that targets all four stages of Vision-Language-Action inference simultaneously and moves end-to-end autonomous driving substantially closer to real-time deployment.
Abstract
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.
FlashVLA is introduced, a streaming action decoding framework that addresses both challenges in a unified formulation and can achieve $\geq$30\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.
Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.
Kemal Oksuz, Alexandru Buburuzan, Yu-Han Yao et al.· 0 citations
A bounded additive pre-softmax attention bias on the visual tokens of detector localized traffic actors on Alpamayo-R1's Qwen3-VL backbone is studied, suggesting the bias governs where the model looks rather than encoding a target behavior.
D. Prasad, Lars Ullrich, Knut Graichen· 1 citation
Gated VLA-Cache is proposed, a lightweight, training-free extension that augments visual-similarity caching with neural introspection that improves reliability when blind caching hurts.
StreamPI is proposed, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters and seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference.
Zhe Liu, Jinghua Hou, Yuxiang Lu et al.· 1 citation· ⚡1
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but...
Yi-Heng Ji, Xing-Ru Zhou, Luis Sentis et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.