Skip to content
#artificial intelligence Book Open access

TileSpMM: A Variable-Size Tiled Algorithm for Sparse Matrix-Matrix Multiplication on Tensor Cores

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 986-996 · 0 citations · 23 references

Abstract

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental primitive in scientific computing and artificial intelligence applications. Modern hardware, notably Tensor Core Units (TCUs), offers immense computational power, creating promising opportunities for SpMM acceleration. However, it is difficult to map sparse matrices with irregular structures onto TCUs due to their requirement for regular operands. Existing fixed-granularity tiling methods frequently face a trade-off between massive zero-padding in sparse regions and poor spatial data locality in dense regions. To bridge this gap, we propose TileSpMM, which breaks the static-granularity bottleneck through a variable-size tiling algorithm that dynamically adapts to local sparsity patterns. Furthermore, TileSpMM is equipped with an adaptive load-balancing strategy and customized granularity-specific kernels to improve hardware utilization and mitigate computation redundancy. Experiments on NVIDIA H100 and RTX 5090 GPUs with a diverse range of benchmark matrices show that TileSpMM delivers overall better performance than existing SpMM methods across the evaluated platforms and datasets. Compared with cuSPARSE, SSpMM, Acc-SpMM and FlashSparse, TileSpMM achieves geometric mean speedups of 4.76 × , 2.74 × , 2.38 × and 1.58 × , respectively.

Read PDF

Similar papers

Book Open access Sep 2026

AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs

AFH-SpMM is a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs that achieves average speedups of 1.33 ×, and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large...

Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia · 0 citations
Open access Aug 2026

DistSpMM: Accelerating Sparse Matrix Dense Matrix Multiplication on GPUs

The proposed DistSpMM proposes DistSpMM, a co-design framework integrating data layout, pipelining, and communication strategies, which introduces HSDMA, a lightweight algorithm that reduces communication by optimizing the dense matrix allocation.

Junyu Gu, Jue Wang, Zhikuang Xin et al. · 0 citations
Book Open access Sep 2026

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

CoTC-SpMM is introduced, a cooperative Tensor–CUDA cores scheme for efficient SpMM on GPUs that first proposes the HTC format to partition sparse matrices into dense and sparse components, enabling specialized kernels to leverage the distinct advantages of different computing units and maximize hardware utilization.

Qi Du, Sheng-Le Lin, Yue-Dan Chen et al. · 0 citations
Open access Sep 2026

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Sheng-Bai Luo, Sheng Ma, Bo Wang et al. · 0 citations
Open access

CB-Sparse:A Cache-Friendly Data Aggregating Algorithm for Block-Based Sparse Matrix Multiplication on GPUs

Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...

Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al. · 0 citations
#machine learning Preprint Aug 2026

SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System

Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy. For the same SpMM on the same matrix, cuSPARSE exhibits a 350x performance gap between CSR and Blocked-ELL. Our study of mult...

Shiyang Li, Yan-Zhi Wang, Ming-Yi Hong · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.