Skip to content

Author

Minyi Guo

We have 5 of 36 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it impose...

Zheng Liu, Zeyu Guo, Zihan Liu et al. · 1 citation
Preprint Sep 2026

Atlas: Algorithm-Hardware Co-Design for On-Device City-Scale 3D Gaussian Splatting in VR

3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, enabling city scale 3DGS on mobile VR devices remains challenging, as the memory requirement of large scale scenes far exceeds the memory capacity of today's mobile GPUs. This paper presents Atlas, an on devic...

He Zhu, Zheng Liu, Xingyang Li et al. · 0 citations
Jul 2026

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations

Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timesteps and self attention computation. As diffusion timesteps are reduced, the computation cost of self attention becomes the dominant bottleneck. Existing acceleration approaches larg...

Wenxuan Miao, Haosong Liu, Weiming Hu et al. · 1 citation
Jul 2026

DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction

DSTAR, a software-hardware co-design framework that accelerates DiT inference by reducing spatial and temporal redundancy and incorporates a sparse attention reuse mechanism to minimize redundant computation in attention layers, and design a specialized hardware accelerator which achieves high efficiency in both latenc...

Chi Zhang, Jieru Zhao, Yu Feng et al. · 2 citations
Preprint Aug 2026

A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

This work proposes FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model, and extends the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping.

Zihan Liu, Jingwen Leng, Yangjie Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.