Jul 2026
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
This paper proposes VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics, and proposes a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation.
Zhuoran Song, Haozhe Jiang, Chunyu Qi et al.
· arXiv.org · 0 citations