Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 5 references
TL;DR
TFS is presented, a tile-aware fusion framework that eliminates this intermediate materialization by fusing SpMM and GeMM directly inside the tile registers of Intel’s Advanced Matrix Extensions (AMX).
Abstract
Graph Neural Network (GNN) inference involves two successive matrix operations per layer: a sparse neighbor aggregation (SpMM) followed by a dense linear transformation (GeMM). The conventional two-step execution materializes a large intermediate matrix Z in main memory, incurring significant memory traffic that dominates end-to-end latency. We present TFS, a tile-aware fusion framework that eliminates this intermediate materialization by fusing SpMM and GeMM directly inside the tile registers of Intel’s Advanced Matrix Extensions (AMX). The key insight is that tile-based sparse computation becomes efficient only when the degree distribution within each tile batch is carefully controlled. TFS introduces degree-aware tile scheduling, which sorts rows by ascending degree so that rows within each 16-row AMX tile batch have similar neighbor counts, raising tile efficiency η to 0.96 on average. Evaluated on 25 sparse matrices from 8+ application domains, including graphs with 65.6 M nodes and 3.6 B edges, TFS achieves up to 6.67 × kernel-level and 4.09 × end-to-end speedup over Intel MKL. An AMX isolation experiment further demonstrates that AMX tiles provide up to 13.7 × speedup over an equivalent AVX-512 BF16 implementation, indicating that tile compute density, not the BF16 format alone, is the primary source of acceleration.
TileGEMM is proposed, a high-performance GEMM implementation on AMX that systematically enhances data reuse across the memory hierarchy, and achieves average speedups of 3.27 × and 1.96 × over AVX-512-based implementations TVM and MKL, respectively.
Kang-Kang Chen, Hua-You Su, Meng-Han Jia et al.· Proceedings of the Internati...· 0 citations
Graph Neural Networks (GNNs) have become a fundamental tool for learning over graph-structured data. Under the message-passing framework, mainstream GNN models alternate between feature transformation and neighborhood aggregation. Fusing these two phases into a node-level pipelined push dataflow, in which each node’s t...
Shi Chen, Jun-Sheng Chang, Yang Guo et al.· ACM Transactions on Architec...· 0 citations
Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...
Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al.· ACM Transactions on Design A...· 0 citations
Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....
Kaiwen Deng, Sifan Sun, Hanjie Liu et al.· IEEE Non-Volatile Memory Sys...· 0 citations
The authors co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live...
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee et al.· 0 citations
Sparse matrix-matrix multiplication (SpMM) is critical for graph analytics and learning tasks, yet its irregular sparsity poses challenges for hardware acceleration. Traditional static dataflows fail to adapt to local sparsity variations, causing load imbalance. We introduce SpMM-GO, a hybrid FPGA accelerator integrati...
Shang-Shang Yao, Yan Yan, Zuo-Ning Chen· IEEE Transactions on Very La...· 0 citations
Related blog posts
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026