Skip to content
Book Open access

Adaptive GPU Sharing for Real-time LLM Serving with Best-effort Workloads

Sep 2026 · Workshop Proceedings of the 55th International Conference on Parallel Processing · 0 citations · 8 references

Abstract

Large Language Models (LLMs) are increasingly deployed in latency-sensitive applications, where real-time serving must satisfy stringent service-level objectives (SLOs). However, request intensities fluctuate over time, and under low load LLM services leave a substantial fraction of GPU task idle. A promising approach to reclaim this capacity is co-locating best-effort workloads with LLM serving via GPU sharing. In this study, we propose AdMix, a scheduling system that enables adaptive GPU sharing between real-time LLM serving and best-effort non-LLM deep learning (DL) inference. AdMix dynamically regulates the best-effort workload’s resource allocation using NVIDIA Multi-Process Service (MPS), guided by real-time monitoring of LLM performance and a lightweight latency estimator. By real-world LLM serving traces, experiments on LLaMA-7B and Qwen-14B paired with representative DL workloads demonstrate that AdMix improves best-effort throughput by 1.3–5 × over MPS baseline while preserving LLM SLO attainment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.