Skip to content

Carryover Drafting: Recycling Rejected States for Speculative Decoding

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

Carryover Drafting is introduced, a parallel draft--verify--draft training that exposes the drafter to inference-aligned rejected states while preserving parallelism across training positions.

Abstract

Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We find that these discarded hidden states generated during target forward retain useful information about future tokens that can improve subsequent drafts. However, realizing this opportunity poses two distinct challenges. At inference, recycling overhead can increase drafting latency, diminishing the speedup gained from increased acceptance length. During training, standard parallel drafter training does not produce inference-aligned rejected states, while obtaining them through sequential rollouts would sacrifice parallelism across training positions. We introduce Carryover Drafting, which addresses both challenges. Carryover recycles rejected target hidden states as temporary KV context, allowing the drafter to selectively attend to them. It reuses the drafter's existing interface and adds only a single learned embedding to distinguish rejected states from committed context. The additional KV context is replaced each drafting round, keeping its length bounded by one proposal block. We introduce parallel draft--verify--draft training that exposes the drafter to inference-aligned rejected states while preserving parallelism across training positions. Experiments with DFlash and a DSpark-derived semi-autoregressive drafter across two target models show that this simple Carryover mechanism improves average acceptance length by 6.5--14.7% and end-to-end vLLM speedup by 7.9--14.4% over the corresponding baselines, with speedup gains reaching 28.8% on translation.

View source

Similar papers

#natural language process... Preprint Sep 2026

DFlow: Enabling Verifier Information Flow in Block Diffusion Speculative Decoding

Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the rejected suffix, preventing the computation spent on...

Yao-Jie Zhang, Lin-Feng Zhang, Bin Cui et al. · 1 citation
#natural language process... Preprint Sep 2026

DEdit: Iterative Draft Editing for Speculative Decoding

Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes...

Long-Xuan Yu, Bingsen Chen, Peng Shi et al. · 0 citations
#natural language process... Preprint Aug 2026

ReTrace: Rejected-Trajectory Conditioning for Speculative Decoding

ReTrace is introduced, a rejected-trajectory conditioning method that conditions each draft block on the rejected suffix from the previous round rather than generating it from fresh mask placeholders alone, indicating that the draft model can retain useful semantic and structural information despite local token-level e...

Luxi Lin, Zhan-Peng Zeng, Shuang Peng et al. · 1 citation
Preprint Sep 2026

NebulaSD: Many-for-Many Speculative Decoding

NebastianSD is presented, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.

Jun-Hao He, Hong-Yang Du · 0 citations
Preprint Aug 2026

DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding

This work proposes a dependent block drafter based on a low-rank latent mixture over token positions, complemented by an acceptance-oriented training objective that directly targets the expected verified length.

Amirmohammad Karimi, Chao Gao, Negar Hassanpour · 1 citation
#natural language process... Preprint Aug 2026

Verification-Aware Training for Speculative Decoding

Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.

Geonmo Gu, Byeongho Heo, Heejae Jun et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.