Aug 2026· ACM Transactions on Architecture and Code Optimization (TACO)· Vol 23, pp. 1 - 25· 0 citations· 44 references
TL;DR
Janus is proposed, a novel resource- and latency-aware operator scheduling framework to enable efficient and practical operator parallelism for DNN inference on GPUs and introduces an effective stream allocation mechanism that fully incorporates operator resource constraints and latency heterogeneity.
Abstract
With the growing deployment of Deep Neural Networks (DNNs) in latency-critical services, optimizing inference efficiency on GPUs has become crucial. While exploiting operator parallelism offers a promising avenue to accelerate inference and improve hardware utilization, existing approaches often overlook two critical factors: hardware resource constraints and latency heterogeneity across operators. This oversight creates a significant discrepancy between the intended schedule and actual runtime behavior, severely degrading inference performance and GPU utilization. To address this, we propose Janus, a novel resource- and latency-aware operator scheduling framework to enable efficient and practical operator parallelism for DNN inference on GPUs. Janus introduces an effective stream allocation mechanism that fully incorporates operator resource constraints and latency heterogeneity. Specifically, it classifies operators into distinct priority levels and leverages the CUDA stream priority mechanism to map them to corresponding streams. By doing so, Janus achieves highly efficient operator parallelism by ensuring that the high-efficiency schedule is faithfully executed at runtime, while simultaneously enabling the hardware scheduler to adaptively scavenge transiently idle resources. We implement a prototype of Janus in PyTorch and conduct comprehensive evaluations using eight representative DNN models on both NVIDIA RTX A5000 and H800 GPUs. Experimental results show that with only a minimal one-time profiling overhead of mere seconds, Janus achieves average inference speedups of 2.13× and 3.41× over PyTorch on the RTX A5000 and H800 GPUs, respectively. Compared with Opara, a state-of-the-art operator-level scheduling framework, Janus further delivers average speedups of 1.20× (RTX A5000) and 1.11× (H800). Moreover, Janus improves GPU utilization. On the RTX A5000, it achieves 1.16× the SM occupancy and 1.10× the SM active rate of Opara on average, which exhibits the highest resource utilization among all baselines.
Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...
Xiaoyang Sun, Jie Xu, Zheng Wang· IEEE Transactions on Paralle...· 0 citations
A time-varying integer program to minimize the long-term total cost of the edge AI inference system, including the inference latency, the inference error rate, the query-dispatching communication cost, and the energy consumption, subject to resource and workload constraints is proposed.
Ming-Tao Ji, Hehan Zhao, Lei Jiao et al.· Science China Information Sc...· 0 citations
ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.
Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al.· 0 citations
A standards-based, multi-tenant cloud inference framework that integrates OpenStack orchestration with Single Root I/O Virtualization (SR-IOV)-enabled graphics processing unit (GPU) partitioning to achieve predictable and isolated real-time inference execution.
Rui Ma, Bing-Feng Shi, Xing-Run Ma et al.· Journal of ICT Standardizati...· 0 citations
Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .
Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al.· 0 citations
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of infe...
Jaehwan Lee, Sang-Min Lee, Chaewon Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.