Skip to content

A dual-stream spectral network for versatile full-band speech enhancement.

Sep 2026 · Journal of the Acoustical Society of America · Vol 160 3, pp. 2272-2287 · 0 citations · 51 references
Medicine

TL;DR

Evaluation results demonstrate that DSVN improves restoration quality under both single- and mixed-degradation scenarios, and shows clear advantages in non-intrusive quality assessment, spectral distance, and automatic speech recognition evaluation, indicating that the enhanced speech is not only cleaner but also more useful for downstream systems.

Abstract

Full-band speech restoration aims to recover high-fidelity speech from signals degraded by noise, reverberation, bandwidth limitation, or their combinations. Existing methods are typically designed for a specific type of degradation, which limits their flexibility when multiple distortions coexist. To address this problem, this paper proposes a dual-stream versatile speech enhancement network (DSVN)-a unified task-conditioned system for full-band speech restoration. The model introduces a dual-stream spectral block to jointly model frequency-wise spectral structures and temporal dynamics. A band-refill decoder is further designed to recover missing or degraded high-frequency components. Task-conditioned modulation is introduced to enable the same network to adapt to different objectives. Evaluation results demonstrate that DSVN improves restoration quality under both single- and mixed-degradation scenarios. The model shows clear advantages in non-intrusive quality assessment, spectral distance, and automatic speech recognition evaluation, indicating that the enhanced speech is not only cleaner but also more useful for downstream systems. Additional analysis of internal responses and ablation studies provides evidence that the proposed modules contribute to different aspects of restoration.

View source

Similar papers

Preprint Aug 2026

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

An Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound and combines interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practic...

Ali Rajabi, Xiang-Wei Zhou · 0 citations
Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Review Open access Sep 2026

Advances in Speech Enhancement: A Comprehensive Review of Noise Suppression Techniques

Over the past several decades, numerous methods have been developed to improve the signal-to-noise ratio, perceptual quality, and intelligibility of speech. In practice, no single method is universally optimal, as each category exhibits distinct strengths and limitations under specific acoustic conditions. The proposed...

Pushpraj Tanwar, A. Somkuwar, Rakesh Kumar Gumasta · 0 citations
Open access Aug 2026

Efficient speech denoising using CleanUNet optimized with Mamba and hybrid spectral loss

Speech enhancement aims to recover clean speech from signals degraded by noise and adverse acoustic conditions, such as background interference and echo. CleanUNet has emerged as an effective solution for causal speech denoising, but its high computational demands limit deployment on resource-constrained devices. In th...

Matheus Vieira da Silva, João Fernando Mari, A. Backes · 0 citations
Preprint Aug 2026

Towards Balanced Spectral Reconstruction: Spectrally Adaptive Loss for Streaming Speech Enhancement

This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency regions caused by the magnitude-phase compensation effect. The proposed sigmoid-weighted loss applies a smooth frequency-dependent modulation t...

Haixin Zhao, Nilesh Madhu · 0 citations
Open access Aug 2026

Single-Channel Speech Enhancement Method Based on Deep Neural Networks

In practical voice communication and processing, speech signals are extremely sensitive to background noise, equipment noise, and interference from complex environments, leading to a decline in speech quality and clarity. This, in turn, affects the overall performance of downstream systems such as speech recognition, s...

Wen Fan, Wei-Yu Liang, Duo-Duo Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.