Skip to content
Preprint

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Aug 2026 · 3 citations · ⚡ 1 influential · 69 references
Computer Science

TL;DR

TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts, and shows that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum.

Abstract

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.

View source

Similar papers

Conference Open access Sep 2026

SplitScaling: Adaptive Scaling for Disaggregated LLM Serving Against Traffic Bursts via DRL

This work proposes a Deep Reinforcement Learning-based auto-scaling framework tailored for the PD architecture that enables the agent to capture non-linear load dynamics, thereby achieving decoupled and precise scaling for prefill and decode pools.

Wei Xiao, Xue-Feng Huang, Wei-Jia Shi et al. · 0 citations
Preprint Aug 2026

Scheduling Mixed RL Rollouts Beyond Prefix Locality

MISA-T, a routing-layer admission policy for mixed rollout serving that combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting, is presented.

Zetao Hong, Song Yuan, Yuan-Hao Ding et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

This work proposes Drift-Aware Sparse Routing (DRS), a nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models.

Cheung-Hao Lee, Patrick Wong · 0 citations
Preprint Aug 2026

P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving

Prefill-Pressure Adaptive Scheduling (P-PAS), a lightweight policy that dynamically adapts the scheduling budget based on concurrent prefill and decode state, is introduced, maintaining low end-to-end latency across changing load regimes, avoiding the limitations of a fixed MBT.

Timo Sämann · 0 citations
#artificial intelligence Preprint Oct 2026

Jumping the Line: Exploiting Length Predictions in LLM Scheduling

Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an...

Yu-Yang Dai, Rana Shahout, Mahmood Sharif · 0 citations
#machine learning Preprint Oct 2026

Flash-OPD: Fast On-Policy Distillation

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision comp...

Wei Chen, Junle Chen, Yi-Tong Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.