Skip to content

Efficient One-to-Many Translation with Joint Multi-Stream Diffusion

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

This work explores how diffusion can enable multilingual translation with a discrete diffusion framework that refines all target languages in parallel, achieving sublinear latency scaling with the number of targets, and supports deployment as a single unified model to replace multiple independent systems.

Abstract

One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number of target languages. We explore how diffusion can enable multilingual translation with a discrete diffusion framework that refines all target languages in parallel, achieving sublinear latency scaling with the number of targets, and supports deployment as a single unified model to replace multiple independent systems. Conditioned on a continuous semantic anchor rather than source tokens, our framework supports zero-shot transfer to unseen source languages without retraining, maintaining approximately $75\%$ of its supervised translation quality on zero-shot sources. We investigate the quality-latency frontier and find that with accelerated sampling, it achieves comparable supervised quality to AR baselines with a $2 \times$ speedup and $11.9\%$ better zero-shot BLEU. These results highlight the potential of joint multi-stream diffusion as a practical and flexible alternative for efficient one-to-many translation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Less Uniform Discrete Diffusion is More Powerful and Scalable

Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM fram...

Kai-Bo Wang, Ding Ding, Fang-Yuan Ding et al. · 0 citations
#artificial intelligence Preprint Sep 2026

One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be diff...

Yu-Lin Yuan, Ying Zhang, Xiang-Ming Meng · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Stream Simultaneous Translation via 2D Grid Attention

A dual-stream attention framework is proposed that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization.

Yu Pu, Wei-Qiang Zhang · 0 citations
#machine learning Preprint Sep 2026

Distribution Matching Distillation for Continuous Diffusion Language Models

Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student'...

Paul Le Van Kiem, Dario Shariatian, Umut Simsekli et al. · 0 citations
#machine learning Preprint Sep 2026

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

Language generation is almost universally treated as a sequential process: autoregressive models emit one token at a time, while diffusion language models replace token-level seriality with a long trajectory of iterative refinement. In this work, we introduce PlaidQ, a 0.7B continuous diffusion language model for code...

Fred Zhangzhi Peng, Kai-Wen Zheng, An-Ru R. Zhang · 1 citation · ⚡1
#natural language process... Preprint Sep 2026

Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR

Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed toward high-resource languages and degrades sharply for languages with limited labeled data and pre-training exposure. To address this, we investigate parameter-efficie...

Thai Thi Thanh Thao Dang, Meng-Jie Qian, Kate Knill · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.