Skip to content
Preprint

FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

Aug 2026 · 4 citations · 26 references
Computer Science

TL;DR

FlashVLA is introduced, a streaming action decoding framework that addresses both challenges in a unified formulation and can achieve $\geq$30\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.

Abstract

Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approaches often fail to jointly achieve low-latency inference and accurate, temporally consistent asynchronous execution. We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challenges in a unified formulation. FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and decodes them using chunk-wise causal attention. This design allows FlashVLA to produce one executable action chunk per inference step. Moreover, its chunk-wise autoregressive formulation implicitly preserves action continuity, enabling smooth asynchronous execution without extra future-state conditioning. Across extensive simulated and real-world experiments, FlashVLA substantially improves inference speed while maintaining strong task performance. It can achieve $\geq$30\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.

View source

Similar papers

Preprint Sep 2026

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respon...

Yu-Han Chen, Ke Yu, Peng-Fei Liu et al. · 0 citations
Preprint Aug 2026

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

FlashDrive is proposed, an algorithm-system co-design framework that targets all four stages of Vision-Language-Action inference simultaneously and moves end-to-end autonomous driving substantially closer to real-time deployment.

Ze-Kai Li, Yihao Liang, Hong-Fei Zhang et al. · 2 citations
Preprint Aug 2026

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI is proposed, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters and seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference.

Zhe Liu, Jinghua Hou, Yuxiang Lu et al. · 1 citation · ⚡1
Preprint Sep 2026

Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but...

Yi-Heng Ji, Xing-Ru Zhou, Luis Sentis et al. · 0 citations
Preprint Sep 2026

Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control

Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physica...

Zi-Bo Wang, Hao-Chen Han, Peng-Zhen Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.