Skip to content

Author

Haibing Guan

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

This paper proposes VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics, and proposes a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation.

Zhuoran Song, Haozhe Jiang, Chunyu Qi et al. · 0 citations
Preprint Aug 2026

Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit t...

Chunyu Qi, Zhuoran Song, Jian Weng et al. · 0 citations
#natural language process... Preprint Sep 2026

PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition

PCoMoE is presented, a path-compositional execution framework that shifts MoE inference from coarse-grained expert selection to fine-grained path composition and achieves up to a 1.31x end-to-end inference speedup while enhancing model accuracy by 10%.

Zi-Yan Gan, Fangxin Liu, Chenyang Guan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.