Skip to content
Open access

Deep learning-based multi-speaker separation and speech enhancement for forensic audio analysis

Aug 2026 · Nature Journal of Emerging Sciences Technologies and Innovations · 0 citations

TL;DR

A hybrid deep learning framework for multi-speaker separation and speech enhancement by integrating Gated Convolutional Neural Networks (GCNNs) for speech source separation with Long Short-Term Memory (LSTM) networks for temporal speech enhancement is proposed.

Abstract

The increasing use of audio recordings in criminal investigations has created a growing demand for intelligent forensic audio analysis systems capable of recovering intelligible speech from acoustically challenging environments. Forensic recordings frequently contain overlapping speakers, background conversations, and environmental noise, making reliable speaker identification and evidence extraction difficult. This study proposed a hybrid deep learning framework for multi-speaker separation and speech enhancement by integrating Gated Convolutional Neural Networks (GCNNs) for speech source separation with Long Short-Term Memory (LSTM) networks for temporal speech enhancement. Unlike conventional approaches that treat speech separation and enhancement as independent tasks, the proposed framework jointly optimizes both processes within a unified architecture while preserving forensic audio integrity. The framework further integrates time-frequency masking, adaptive Wiener filtering, and spectral gain enhancement to suppress background noise while preserving speech fidelity and enhancing low-level background speech that may contain valuable forensic information. A custom dataset comprising 500 two-speaker conversations was developed and evaluated across diverse acoustic environments with signal-to-noise ratios ranging from −20 dB to +20 dB. Experimental results demonstrated an average Signal-to-Distortion Ratio (SDR) improvement of 9.4 dB, an average Signal-to-Noise Ratio (SNR) improvement of 5.2 dB, and a Diarization Error Rate (DER) of 6.4%. The framework also achieved an average processing latency of 1.2 seconds for a 20-second audio segment, indicating its suitability for near real-time forensic applications. The proposed framework substantially improves signal quality, speech intelligibility and speaker discrimination while preserving evidential integrity, thereby providing an effective decision-support tool for forensic audio analysis in criminal investigations.

Read PDF

Similar papers

Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Preprint Aug 2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual ma...

Yan-Qiu Li, Yang Xiao, Jisheng Bai et al. · 0 citations
Preprint Aug 2026

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

An Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound and combines interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practic...

Ali Rajabi, Xiang-Wei Zhou · 0 citations
Preprint Sep 2026

Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized...

Rayhan Rashed, Senja Filipi, Ross Cutler · 0 citations
Open access 2026

A Multi-Domain Feature Framework for Robust Deepfake Audio Detection

The results indicate that multi-domain feature fusion offers a practical and generalizable solution for real-world deepfake audio detection, particularly in environments involving diverse codecs and synthesis techniques.

Akshat Chhatriwala, Ishita Akolkar, Namrata Shroff et al. · 0 citations
Preprint Aug 2026

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

DAVE is presented, a decoupled audio-visual enhancement framework for real-world speech separation that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics.

Wei Zhou, Wan-Yi Ning, Yi Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.