Skip to content
Preprint

Deep Neural Compression for RIR-Characterized Acoustic Environments with Structure-Aware Constraints

Sep 2026 · 0 citations · 24 references
Engineering

TL;DR

An EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels, which achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.

Abstract

Room impulse responses (RIRs) characterize the acoustic environment of a room by capturing how sound propagates and decays within an enclosed space. In applications such as immersive audio rendering, accurate acoustic reconstruction often relies on spatially densely sampled RIRs. This consequently gives rise to a large volume of RIR data, imposing a substantial burden on storage. Although recent neural audio codecs provide an effective framework for low-bitrate compression, their training objectives are mainly tailored to speech and general audio, and are therefore not well aligned with the acoustic characteristics of RIRs. Therefore, we propose an EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels. Specifically, at the RIR level, structure-aware constraints are imposed on the global decay behavior and local energy distribution of RIRs through energy decay curve (EDC) regularization and a short-time window energy constraint, while at the reverberant-speech level, reverberant-speech supervision is further introduced to constrain the consistency of the reverberant speech generated by the reconstructed RIRs. Experimental results show that, at a low bitrate of 375 bps, the proposed method achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.

View source

Similar papers

Preprint Aug 2026

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

Increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments and shows that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.

Alessia Milo, G. Götz, S. Guðjónsson et al. · 2 citations
Preprint Sep 2026

DiffVQE2: An Efficient Low-delay Diffusion Model for Acoustic Echo and Noise Control

Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods ha...

Haljan Lugo, Ernst Seidel, Pejman Mowlaee et al. · 0 citations
Preprint Sep 2026

Spatial Audio Coding Through Relative Room Impulse Response Estimation

Immersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networ...

N. Bouayed, Adrien Llave, J. Daniel et al. · 0 citations
Conference Open access Sep 2026

BEAT2AASIST: BEATs Feature Splitting with Dual-Branch AASIST for Environmental Sound Deepfake Detection

Recent advances in text-to-audio (TTA) and audio-to-audio (ATA) generation models have enabled the creation of highly realistic environmental sounds, raising growing concerns about malicious audio manipulation in real-world scenarios. To address this emerging threat, the ESDD 2026 Challenge was introduced as the first...

Sanghyeok Chung, Seungsang Oh, Donggun Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs

Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whisper-Tiny and Whisper-Small with EnCodec...

Song-ha Jo, Sehyun Lee, Soyoon Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.