Skip to content

From acoustic scene to sense of place: Sound scene-to-description, a semantic translation engine for sound environment.

Sep 2026 · Journal of the Acoustical Society of America · Vol 160 3, pp. 2301-2314 · 0 citations · 34 references
Medicine

Abstract

Sound environments play a crucial role in human experience, shaping memory, comfort, and sense of place. They provide essential cues for judging safety and atmosphere in both real and virtual settings. Despite the growing availability of large audio libraries, extracting meaningful information from complex and overlapping sound scenes remains difficult. Audio captioning addresses this challenge by translating acoustic scenes into text, yet traditional approaches face clear limitations. Manual annotation is often subjective and incomplete, while sound event detection reduces audio to isolated tags without capturing context, relationships, or temporal dynamics. To overcome these barriers, this study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model. This allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving beyond simple event lists. In both qualitative and quantitative evaluations, SS2D has demonstrated significantly better performance than the audio event method and the image caption method. In terms of practical significance, SS2D eliminates the need for extensive manual labeling and has broad applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.