Skip to content

Author

Shigang Li

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

MX-KMeans: Accelerating K-Means Clustering through Microscaling Quantization

K-Means clustering is a classical unsupervised learning method widely used for its simplicity, efficiency, and broad applicability. In this work, we first analyze the numerical distributions of representative K-Means datasets and identify an opportunity for low-precision acceleration through hardware-native microscalin...

Rong-Tian Fu, Dong-Bo Lv, Xue-Ying Wang et al. · 0 citations
#data science Nov 2026

UltraGNN: A Sparse-Operator-Aware Framework for Accelerating Graph Neural Networks on Tensor Cores

Graph Neural Networks (GNNs) have achieved widespread success from social networks to AI-for-Science. Most existing GNN frameworks adopt scatter-first (edge-centric) or gather-first (vertex-centric) scheduling paradigms for message passing. However, these paradigms are closely tied to traditional CUDA-core execution mo...

Jin-Liang Shi, Shi-Gang Li, Rong-Tian Fu et al. · 0 citations
Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

OmniPipe is proposed, a flexible bidirectional multi-pipeline parallelism scheme for unified dense and MoE LLM training that minimizes the pipeline bubble ratio while effectively overlapping EP communication with computation, enabled by the flexible and scalable parallelism scheme of bidirectional pipelines.

Jun Li, Zhi Ma, Shi-Gang Li · 0 citations
Book Open access Sep 2026

MX-KMeans: Accelerating K-Means Clustering through Microscaling Quantization

K-Means clustering is a classical unsupervised learning method widely used for its simplicity, efficiency, and broad applicability. In this work, we first analyze the numerical distributions of representative K-Means datasets and identify an opportunity for low-precision acceleration through hardware-native microscalin...

Rong-Tian Fu, Dong-Bo Lv, Xue-Ying Wang et al. · 0 citations
#large language models Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

Mixture-of-Experts (MoE) has become the de facto architecture for scaling large language models, offering expanded capacity with manageable compute. Pipeline parallelism (PP) is indispensable for distributed MoE training, but state-of-the-art PP schemes face three major limitations: large pipeline bubbles, high per-sta...

Jun Li, Zhi Ma, Shigang Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.