Skip to content
Book Open access

Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDC

Sep 2026 · Proceedings of the ACM SIGOPS 32nd Symposium on Operating Systems Principles · 0 citations · 43 references

Abstract

Serving offline large language model (LLM) inference workloads (e.g., log summarization and bulk translation) can consume up to 30% of GPUs in production. Despite this significant share, the characteristics of offline inference remain largely understudied. In this paper, we start by analyzing 1.5 million tasks comprising 23 billion requests across text and multi-modal models. We discover that the key properties of offline workloads, namely inherent determinism and throughput orientation, are neither exploited by online LLM serving systems nor by existing offline serving frameworks. We design ACDC, a task-aware offline serving system that leverages these properties along three orthogonal axes, namely resource orchestration, cache management, and execution flow. ACDC has been deployed in production for over three months, achieving up to 3.9× higher output cost-efficiency. Beyond isolated deployment, we further identify that online prefill stages leave up to 38.3% of capacity idle as resource bubbles. We design a hybrid-serving mechanism that harvests these bubbles for offline tasks, improving end-to-end throughput by up to 35.7% over the state-of-the-art alternative.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.