Skip to content
Book Open access

CoTC-SpMM: A Cooperative Tensor–CUDA Cores Scheme for Efficient Sparse Matrix Multiplication

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 29 references

TL;DR

CoTC-SpMM is introduced, a cooperative Tensor–CUDA cores scheme for efficient SpMM on GPUs that first proposes the HTC format to partition sparse matrices into dense and sparse components, enabling specialized kernels to leverage the distinct advantages of different computing units and maximize hardware utilization.

Abstract

Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in scientific computing and artificial intelligence. However, the inherent sparsity and irregularity of real-world datasets present significant challenges for developing high-performance SpMM kernels on modern GPUs. Existing approaches typically focus on a single type of compute unit, leaving the potential for cooperative parallelism between heterogeneous GPU cores largely under-explored. This paper introduces CoTC-SpMM, a cooperative Tensor–CUDA cores scheme for efficient SpMM on GPUs. Specifically, we first propose the HTC format to partition sparse matrices into dense and sparse components, enabling specialized kernels to leverage the distinct advantages of different computing units and maximize hardware utilization. Moreover, we implement several low-level runtime optimizations, including 1-D resource mapping for load balancing, software pipelining to hide memory latency, and PTX-level instruction tuning to enhance SpMM throughput. Experimental results on NVIDIA A100 and H800 GPUs demonstrate that CoTC-SpMM achieves substantial performance speedups over state-of-the-art implementations.

Read PDF

Similar papers

Book Open access Sep 2026

AFH-SpMM: Auto-Fit Heterogeneous Block Sparse-Dense Matrix Multiplication on Tensor Core GPUs

AFH-SpMM is a novel Auto-Fit Heterogeneous SpMM framework designed for adaptively parallelizing sparse-dense matrix multiplication on Tensor Core-equipped GPUs that achieves average speedups of 1.33 ×, and often leads cuSPARSE, ASpT, Sputnik, RoDe, Acc-SpMM, and MP-SpMM, with especially strong gains on medium and large...

Zhi-Rui Chen, Heng Zhang, Kai-Fan Jia · 0 citations
Open access Sep 2026

ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging

This work proposes the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support the dataflow paradigm while enhancing data reuse, and decouples the multiplication and merging phases.

Sheng-Bai Luo, Sheng Ma, Bo Wang et al. · 0 citations
Book Open access Sep 2026

BAG: Faster Matrix Multiplication on a Single GPU

BAG (Basis Alternative Matrix Multiplication on GPUs), a GPU-oriented implementation of ABMM for the NVIDIA Ampere architecture is presented, designed to shrink workspace and eliminate redundant global-memory traffic, and introduce a cost-model-based recursion policy together with a Roofline-guided blocking strategy to...

Yao Liu, Ye-Wen Li, Zhong-Hai Zhang et al. · 0 citations
Open access

CB-Sparse:A Cache-Friendly Data Aggregating Algorithm for Block-Based Sparse Matrix Multiplication on GPUs

Sparse matrix multiplications—including SpMV, SpMM, and SpGEMM—are fundamental to scientific computing, graph analytics, and machine learning. Despite extensive GPU-focused optimizations such as custom sparse formats and load balance, CSR-style and block-based methods can still underexploit fine-grained cache locality...

Xing Cong, Fu-Kai Sun, Yi-Ding Liu et al. · 0 citations
Book Open access Aug 2026

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

DB-SpMSpV is presented, a dual-view blocked SpMSpV framework for dynamic GPU workloads that uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance and is integrated into DB-BFS and DB-Decoding.

Xing Cong, Chen-Hao Xie, Rui Wang et al. · 0 citations
Preprint Sep 2026

Rethinking Sparse Formats for RISC-V: A Hierarchical Approach to High-Performance SpMV

The sparse matrix-vector multiplication (SpMV) algorithm is a fundamental computational kernel of linear algebra and serves as a building block for numerous applications, primarily iterative solvers for systems of linear equations used in scientific and engineering simulations. This paper compares vectorized implementa...

A. Pirova, Anastasia Vodeneeva, K. Kovalev et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.