Book
Open access
Jul 2026
Toward Low-Latency and Memory-Efficient Deployment of Irregular Sparse Deep Learning Workloads
This work introduces a tile-aware scheduling framework for efficient sparse Vision Transformer execution on GPUs and introduces a training-aware extension that reuses the inference tile schedule and augments it with backward computation and activation-memory strategies.
Changxin Li
· IEEE International Symposium... · 0 citations