ExplorerRoboticsRobotics
Research PaperResearchia:202608.28012

FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

Zekai Li

Abstract

Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approach...

Submitted: August 28, 2026Subjects: Robotics; Robotics

Description / Details

Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approaches often fail to jointly achieve low-latency inference and accurate, temporally consistent asynchronous execution. We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challenges in a unified formulation. FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and decodes them using chunk-wise causal attention. This design allows FlashVLA to produce one executable action chunk per inference step. Moreover, its chunk-wise autoregressive formulation implicitly preserves action continuity, enabling smooth asynchronous execution without extra future-state conditioning. Across extensive simulated and real-world experiments, FlashVLA substantially improves inference speed while maintaining strong task performance. It can achieve \geq30,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.


Source: arXiv:2608.27384v1 - http://arxiv.org/abs/2608.27384v1 PDF: https://arxiv.org/pdf/2608.27384v1 Original Link: http://arxiv.org/abs/2608.27384v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 28, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark