Skip to content
Open access

Janus: Realizing Practical Operator Parallelism for Latency-Sensitive DNN Inference on GPUs

Aug 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · Vol 23, pp. 1 - 25 · 0 citations · 44 references

TL;DR

Janus is proposed, a novel resource- and latency-aware operator scheduling framework to enable efficient and practical operator parallelism for DNN inference on GPUs and introduces an effective stream allocation mechanism that fully incorporates operator resource constraints and latency heterogeneity.

Abstract

With the growing deployment of Deep Neural Networks (DNNs) in latency-critical services, optimizing inference efficiency on GPUs has become crucial. While exploiting operator parallelism offers a promising avenue to accelerate inference and improve hardware utilization, existing approaches often overlook two critical factors: hardware resource constraints and latency heterogeneity across operators. This oversight creates a significant discrepancy between the intended schedule and actual runtime behavior, severely degrading inference performance and GPU utilization. To address this, we propose Janus, a novel resource- and latency-aware operator scheduling framework to enable efficient and practical operator parallelism for DNN inference on GPUs. Janus introduces an effective stream allocation mechanism that fully incorporates operator resource constraints and latency heterogeneity. Specifically, it classifies operators into distinct priority levels and leverages the CUDA stream priority mechanism to map them to corresponding streams. By doing so, Janus achieves highly efficient operator parallelism by ensuring that the high-efficiency schedule is faithfully executed at runtime, while simultaneously enabling the hardware scheduler to adaptively scavenge transiently idle resources. We implement a prototype of Janus in PyTorch and conduct comprehensive evaluations using eight representative DNN models on both NVIDIA RTX A5000 and H800 GPUs. Experimental results show that with only a minimal one-time profiling overhead of mere seconds, Janus achieves average inference speedups of 2.13× and 3.41× over PyTorch on the RTX A5000 and H800 GPUs, respectively. Compared with Opara, a state-of-the-art operator-level scheduling framework, Janus further delivers average speedups of 1.20× (RTX A5000) and 1.11× (H800). Moreover, Janus improves GPU utilization. On the RTX A5000, it achieves 1.16× the SM occupancy and 1.10× the SM active rate of Opara on average, which exhibits the highest resource utilization among all baselines.

Read PDF

Similar papers

Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
Aug 2026

Learning to schedule: interference-aware optimization for edge AI inference on shared GPUs

A time-varying integer program to minimize the long-term total cost of the edge AI inference system, including the inference latency, the inference error rate, the query-dispatching communication cost, and the energy consumption, subject to resource and workload constraints is proposed.

Ming-Tao Ji, Hehan Zhao, Lei Jiao et al. · 0 citations
Preprint Aug 2026

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters

ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.

Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al. · 0 citations
Open access Aug 2026

Design of a Multi-Tenant Real-Time Inference Framework Based on OpenStack and SR-IOV GPU Virtualization

A standards-based, multi-tenant cloud inference framework that integrates OpenStack orchestration with Single Root I/O Virtualization (SR-IOV)-enabled graphics processing unit (GPU) partitioning to achieve predictable and isolated real-time inference execution.

Rui Ma, Bing-Feng Shi, Xing-Run Ma et al. · 0 citations

Procyon: Promoting Fine-Grain Multi-Tenancy to Optimize Sparse Streaming Accelerators

Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .

Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al. · 0 citations
#machine learning Preprint Sep 2026

Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of infe...

Jaehwan Lee, Sang-Min Lee, Chaewon Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.