One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be diff...