Skip to content

TFS: Tile-Aware SpMM–GeMM Fusion for Accelerating GNN Inference on Intel AMX

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 5 references

TL;DR

TFS is presented, a tile-aware fusion framework that eliminates this intermediate materialization by fusing SpMM and GeMM directly inside the tile registers of Intel’s Advanced Matrix Extensions (AMX).

Abstract

Graph Neural Network (GNN) inference involves two successive matrix operations per layer: a sparse neighbor aggregation (SpMM) followed by a dense linear transformation (GeMM). The conventional two-step execution materializes a large intermediate matrix Z in main memory, incurring significant memory traffic that dominates end-to-end latency. We present TFS, a tile-aware fusion framework that eliminates this intermediate materialization by fusing SpMM and GeMM directly inside the tile registers of Intel’s Advanced Matrix Extensions (AMX). The key insight is that tile-based sparse computation becomes efficient only when the degree distribution within each tile batch is carefully controlled. TFS introduces degree-aware tile scheduling, which sorts rows by ascending degree so that rows within each 16-row AMX tile batch have similar neighbor counts, raising tile efficiency η to 0.96 on average. Evaluated on 25 sparse matrices from 8+ application domains, including graphs with 65.6 M nodes and 3.6 B edges, TFS achieves up to 6.67 × kernel-level and 4.09 × end-to-end speedup over Intel MKL. An AMX isolation experiment further demonstrates that AMX tiles provide up to 13.7 × speedup over an equivalent AVX-512 BF16 implementation, indicating that tile compute density, not the BF16 format alone, is the primary source of acceleration.

Read PDF

Similar papers

Book Open access Sep 2026

TileGEMM: Boosting the Performance of GEMM on AMX-Powered CPUs by Exploiting Data Reuse

TileGEMM is proposed, a high-performance GEMM implementation on AMX that systematically enhances data reuse across the memory hierarchy, and achieves average speedups of 3.27 × and 1.96 × over AVX-512-based implementations TVM and MKL, respectively.

Kang-Kang Chen, Hua-You Su, Meng-Han Jia et al. · 0 citations
#graph neural networks Open access Sep 2026

PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution

Graph Neural Networks (GNNs) have become a fundamental tool for learning over graph-structured data. Under the message-passing framework, mainstream GNN models alternate between feature transformation and neighborhood aggregation. Fusing these two phases into a node-level pipelined push dataflow, in which each node’s t...

Shi Chen, Jun-Sheng Chang, Yang Guo et al. · 0 citations
Open access Sep 2026

DCSR-GCN: A High-Performance GCN Accelerator Based on Dynamic Compression and Sparsity Reordering

Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...

Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al. · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
Preprint Sep 2026

Hardware Acceleration of Block-Diffusion LLM for Edge Devices

The authors co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live...

Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee et al. · 0 citations
Oct 2026

SpMM-GO: Sparse Matrix-Matrix Multiplication Acceleration via Hybrid Gustavson and Outer-Product Dataflows

Sparse matrix-matrix multiplication (SpMM) is critical for graph analytics and learning tasks, yet its irregular sparsity poses challenges for hardware acceleration. Traditional static dataflows fail to adapt to local sparsity variations, causing load imbalance. We introduce SpMM-GO, a hybrid FPGA accelerator integrati...

Shang-Shang Yao, Yan Yan, Zuo-Ning Chen · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.