Skip to content

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Jul 2026 · arXiv.org · Vol abs/2607.24148 · 0 citations · 70 references
Computer Science

TL;DR

This paper proposes VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics, and proposes a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation.

Abstract

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.

View source

Similar papers

Preprint Sep 2026

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...

Rithvik Jonna, Man Namgung, Aakash Gurram et al. · 0 citations
Preprint Sep 2026

P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration

Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence leng...

Hai-Zhao Jing, Zhen-Hao Shang, Hao-Kui Zhang et al. · 0 citations
Preprint Aug 2026

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it impose...

Zheng Liu, Zeyu Guo, Zihan Liu et al. · 1 citation
Open access Sep 2026

Hardware-Aware Acceleration of Open-Vocabulary Multi-Object Navigation on Edge GPUs

Persistent open-vocabulary navigation enables robots to search for sequential language-specified objects using reusable visual semantic evidence. On edge GPUs, the pipeline is constrained by foundation-model inference, semantic projection, persistent map movement, frontier processing, and repeated target detection. Usi...

M. Akor, Heoncheol Lee · 0 citations
Book Open access Aug 2026

MVP: A Mobile 3D-Stacked VLM Accelerator for Efficient Video Understanding by Leveraging Dynamic Sparse Attention Patterns

MVP, a 3D-stacked VLM accelerator featuring context-aware sparse attention (CASA) and online workload-aware hybrid parallelism scheduling, is introduced, which prunes redundant attention computation FLOPS and adaptively balances computation across hybrid bonding (HB) based many-core NoC architecture.

Yifan Ding, Qianxu Wang, Dunshan Yu et al. · 0 citations
Preprint Aug 2026

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

This work presents a five-step methodology for zero GPU fallback DLA INT8 deployment of classification backbones, comprising architecture adaptation, manual dynamic range workaround to rescue TensorRT's implicit quantization, and generalizes to any detection-classification edge pipeline.

V. Raju · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.