Skip to content
Preprint

NebulaSD: Many-for-Many Speculative Decoding

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

NebastianSD is presented, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.

Abstract

Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Drafting and verification, however, exhibit different service characteristics and favor different batch configurations, making fixed draft-target coupling inefficient under concurrent workloads. Existing distributed designs can physically separate the two stages, but often retain request or batch affinities that prevent their capacities from being shared globally. We present NebulaSD, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools. Such dynamic reassignment removes fixed worker locality, requiring request states to be made available at newly selected workers without introducing migration stalls. NebulaSD addresses this challenge through worker-triggered batch reconstruction and asynchronous KV-state preparation overlapped with model execution. We evaluate NebulaSD from both system and scaling perspectives, showing that dynamic pooling improves request-round processing rate by 50.4% over a physically disaggregated baseline and 72.6% over co-located execution on a four-GPU deployment while substantially increasing effective GPU utilization. Profile-driven simulations further show approximately proportional compute-side capacity scaling under idealized state movement.

View source

Similar papers

#machine learning Preprint Sep 2026

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

This work proposes ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components, which achieves speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\% over the strongest prior self-speculative baselines.

Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is...

Hao-Yuan He, Peng-Fei Liu, Si-Shi Shen et al. · 0 citations
Preprint Sep 2026

Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference

LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs. Production systems, including vLLM's batch-invariant mode and SGLang's deterministic mode, target this goal but lack a formal...

Jian-Xing Qin, Alexander Du, Dan-Feng Zhang et al. · 0 citations
Preprint Aug 2026

CoRun: Padding is Simple and Efficient for Deterministic LLM Inference

CoRun is presented, a scheduling-based system that achieves deterministic inference without requiring batch invariance, and employs isolated prefill and fixed-shape batched decode to handle the two stages of LLM inference, respectively, leveraging CUDA graphs for efficient execution and simplified implementation.

Shiju Zhao, Jiacheng Yang, Qi-Hang Chen et al. · 0 citations
#machine learning Preprint Sep 2026

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Be...

Jungseob Lee, Sugyeong Eo · 0 citations
#artificial intelligence Preprint Sep 2026

Tsubame: Tree Replay for Diffusion-Based Speculative Decoding

Context-aware dynamic trees allocate the speculative decoding budget according to draft path probabilities, adapting their depth and branching to the current context. Under stochastic decoding, however, we find that this structural advantage does not always compensate for the acceptance gains of random sampling paired...

Yepeng Weng, Qiao Hu, T. Yairi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.