This paper proposes VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics, and proposes a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation.
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit t...
Chunyu Qi, Zhuoran Song, Jian Weng et al.· 0 citations
PCoMoE is presented, a path-compositional execution framework that shifts MoE inference from coarse-grained expert selection to fine-grained path composition and achieves up to a 1.31x end-to-end inference speedup while enhancing model accuracy by 10%.
Zi-Yan Gan, Fangxin Liu, Chenyang Guan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.