Skip to content
Preprint

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

DAVE is presented, a decoupled audio-visual enhancement framework for real-world speech separation that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics.

Abstract

Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions. Existing approaches usually fuse visual features directly into the separation network, making them vulnerable to degraded visual signals. In this paper, we present DAVE, a decoupled audio-visual enhancement framework for real-world speech separation. Firstly, to address the data scarcity issue, we construct DAVE-Corpus, a large-scale training corpus with 219,411 mixtures generated from public meeting corpora through combinatorial acoustic augmentation. Then, we introduce a progressive multi-objective optimization strategy to jointly improve speech separation, intelligibility, speaker identity preservation, and perceptual quality. We further develop a certified selective enhancement chain that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. Experimental results on the Real-World Audio-Visual Speech Enhancement Challenge demonstrate the robustness of DAVE under both real-world mixed scenarios and visual degradation conditions.

View source

Similar papers

Preprint Aug 2026

Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement

Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. T...

Tongtao Ling, Zhong-Qiu Wang · 0 citations
Preprint Aug 2026

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video

Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-...

Kai Li, Wen-Ze Ren, Jun-Jie Li et al. · 0 citations
Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Preprint Aug 2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual ma...

Yan-Qiu Li, Yang Xiao, Jisheng Bai et al. · 0 citations
Preprint Sep 2026

Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized...

Rayhan Rashed, Senja Filipi, Ross Cutler · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.