Skip to content

From acoustic scene to sense of place: Sound scene-to-description, a semantic translation engine for sound environment.

Sep 2026 · Journal of the Acoustical Society of America · Vol 160 3, pp. 2301-2314 · 0 citations · 34 references
Medicine

TL;DR

This study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model, which allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving beyond simple event lists.

Abstract

Sound environments play a crucial role in human experience, shaping memory, comfort, and sense of place. They provide essential cues for judging safety and atmosphere in both real and virtual settings. Despite the growing availability of large audio libraries, extracting meaningful information from complex and overlapping sound scenes remains difficult. Audio captioning addresses this challenge by translating acoustic scenes into text, yet traditional approaches face clear limitations. Manual annotation is often subjective and incomplete, while sound event detection reduces audio to isolated tags without capturing context, relationships, or temporal dynamics. To overcome these barriers, this study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model. This allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving beyond simple event lists. In both qualitative and quantitative evaluations, SS2D has demonstrated significantly better performance than the audio event method and the image caption method. In terms of practical significance, SS2D eliminates the need for extensive manual labeling and has broad applications.

View source

Similar papers

Preprint Aug 2026

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion r...

Zhi Zeng, Cheng Zhang, Ze-Sheng Yang et al. · 2 citations
#machine learning Preprint Sep 2026

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

STAG is introduced, to the authors' knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs, and provides behavioral support for the faithfulness and selectivity of the explanations.

Lucia Cascone, V. Fraenza, Michele Nappi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

A pretrained text-to audiovisual generation model is adapted through source-conditioned feature modulation to jointly learn object addition and removal, and quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition.

Wei-Han Xu, Kan-Jen Cheng, Koichi Saito et al. · 0 citations
Preprint Sep 2026

Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions

Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a b...

Michel Olvera, Paraskevas Stamatiadis, Chang-Hong Wang et al. · 0 citations
Preprint Sep 2026

Do Audio Representations Compose Additively?

Compositionality, the ability to represent complex acoustic scenes as combinations of simpler sound sources, is central to auditory perception and classical additive signal models. Still, it remains unclear whether modern pre-trained audio representations internalize additive structure without compositional supervision...

Chen-Hao Xue, Zhi-Jin Guo, Joyraj Chakraborty et al. · 0 citations
Preprint Sep 2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...

Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.