Janus is proposed, a novel resource- and latency-aware operator scheduling framework to enable efficient and practical operator parallelism for DNN inference on GPUs and introduces an effective stream allocation mechanism that fully incorporates operator resource constraints and latency heterogeneity.
Yi-Feng Zhang, Hao-Xuan Ma, Yu-Xing Long et al.· ACM Transactions on Architec...· 0 citations
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...
Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al.· 0 citations
With the growing deployment of Deep Neural Networks (DNNs) in latency-critical services, optimizing inference efficiency on GPUs has become crucial. While exploiting operator parallelism offers a promising avenue to accelerate inference and improve hardware utilization, existing approaches often overlook two critical f...
Yifeng Zhang, Haoxuan Ma, Yuxing Long et al.· ACM Transactions on Architec...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.