Graph Neural Network (GNN) inference involves two successive matrix operations per layer: a sparse neighbor aggregation (SpMM) followed by a dense linear transformation (GeMM). The conventional two-step execution materializes a large intermediate matrix Z in main memory, incurring significant memory traffic that domina...
Xiao Yan, Haodong Bian, Xiao-Ying Wang et al.· Proceedings of the Internati...· 0 citations
TFS is presented, a tile-aware fusion framework that eliminates this intermediate materialization by fusing SpMM and GeMM directly inside the tile registers of Intel’s Advanced Matrix Extensions (AMX).
Xiao Yan, Haodong Bian, Xiao-Ying Wang et al.· Proceedings of the Internati...· 0 citations
A compact classifier is trained to rank omitted sentences by whether they are needed to interpret retained text and reinsert the top-ranked candidates without support annotations at inference, training a compact classifier to optimize both relevance and referential completeness.
Zheng Hu, Kai Li, Dapeng Fu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.