Skip to content
Preprint

Dynamic Multi-Byte Prediction With Hierarchical Language Models

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

It is shown that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.

Abstract

Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, we introduce multi-byte prediction (MBP), which generates multiple bytes in parallel, speeding up inference with minimal performance impact and no additional parameters. MBP builds on the popular multi-token prediction (MTP) paradigm with two crucial innovations. First, we introduce a variable-length prediction window that aligns with the latent tokens, or segments, of a hierarchical LM. Second, we implement a novel attention-masking scheme that enables parallel byte prediction without violating causality. We show that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.

View source

Similar papers

#natural language process... Preprint Aug 2026

Toppling the Hierarchy in Byte-level Language Modeling

This work examines recent byte-level models and their failure to perfectly manipulate characters, finding that the hierarchical design itself limits character-level understanding, with pure byte-level models consistently outperforming hierarchical variants on character manipulation tasks.

Lukas Edman, Alexander Fraser · 0 citations
Preprint Aug 2026

Can Large Language Models"Hyper-Thread"?

Large language models generate tokens sequentially, but can they execute multiple tasks concurrently while forming each token? Broader attention allocation may provide a mechanism for such task concurrency. Existing approaches to scaling inference primarily rely on longer generations, more samples, or additional verifi...

Fei Ding · 0 citations
Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.

Clara Meister · 1 citation
#natural language process... Preprint Sep 2026

Line-Coupled Language Model

The Line-Coupled Language Model (LCLM), an autoregressive model that advances multiple text lines together by predicting the next token for every active line while coupling the lines through shared causal context is introduced.

Shi-Yuan Li, Shao-Rong Zhang, Zhaorui Yang et al. · 0 citations
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...

Clara Meister · 0 citations
#artificial intelligence Preprint Sep 2026

Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference

Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deplo...

Suwesh Prasad Sah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.