A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
This paper proposes VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics, and proposes a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation.