Preprint
Aug 2026
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
SIEDD, a discrete diffusion framework for text-guided speech inpainting and editing over hierarchical codec tokens, is introduced and results demonstrate that explicitly modeling the codec hierarchy substantially improves context-preserving speech reconstruction and editing.
Iftach Shoham, Tali Dror, Oren Gal et al.
· 0 citations