From acoustic scene to sense of place: Sound scene-to-description, a semantic translation engine for sound environment.
This study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model, which allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving...