Open access
Aug 2026
SAI: Virtualizing Shared Memory of GPU for AI workload acceleration
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preserving L2 cache performance.
Hanqing Li, Tiejun Li, Sheng Ma et al.
· ACM Transactions on Design A... · 0 citations