UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding
Speculative decoding accelerates language model inference by verifying multiple draft tokens in a single target-model pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, but their effectiveness deteriorates as the entropy of target distributions increases due to insuffic...