Skip to content
Preprint

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

Aug 2026 · 2 citations · 33 references
Computer Science

TL;DR

ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

Abstract

Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

View source

Similar papers

Sep 2026

From acoustic scene to sense of place: Sound scene-to-description, a semantic translation engine for sound environment.

Sound environments play a crucial role in human experience, shaping memory, comfort, and sense of place. They provide essential cues for judging safety and atmosphere in both real and virtual settings. Despite the growing availability of large audio libraries, extracting meaningful information from complex and overlapping sound scenes remains difficult. Audio captioning addresses this challenge by translating acoustic scenes into text, yet traditional approaches face clear limitations. Manual annotation is often subjective and incomplete, while sound event detection reduces audio to isolated tags without capturing context, relationships, or temporal dynamics. To overcome these barriers, this study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model. This allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving beyond simple event lists. In both qualitative and quantitative evaluations, SS2D has demonstrated significantly better performance than the audio event method and the image caption method. In terms of practical significance, SS2D eliminates the need for extensive manual labeling and has broad applications.

Yong-Gai Zhuang, Teng Fei, Yun-Yan Du et al. · 0 citations
Preprint Aug 2026

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8\%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.

Leiye Liu, Miao Zhang, Jia-Hong Jiang et al. · 0 citations
#artificial intelligence Open access Aug 2026

Towards the Vision-Sound-Language-Action paradigm: The HEAR framework for sound-centric manipulation

Humans and animals use sound as a crucial cue for interacting with the physical world, as acoustic events can reveal contact, completion, hidden contents, or process state. Embodied agents should similarly benefit from auditory awareness during manipulation, yet existing Vision-Language-Action (VLA) policies typically rely on persistent visual observations, while audio-aware variants often treat audio as speech, waveform renderings, or fixed preexecution context. Such interfaces can miss transient sounds, such as beeps, clicks, rattles, or collision cues, especially under system latency and open-loop action chunking. We formalize this timing failure as the Blind Execution Interval (BEI), in which critical acoustic evidence may occur after an action chunk begins but disappear before the next policy update. To address this challenge, we introduce Vision-Sound-Language-Action (VSLA), a continuous control paradigm conditioned on vision, streaming audio, language, and proprioception under delayed decision loops. We further present HEAR, a VSLA framework that preserves causal auditory context across execution gaps, performs multimodal reasoning, models near-future audio dynamics during training, and generates smooth action chunks for closed-loop manipulation. To support learning and evaluation, we introduce OpenX-Sound for robotics-specific audio-visual-action pretraining and HEAR-Bench, a benchmark for sound-centric manipulation with strict causal timing constraints. On HEAR-Bench, HEAR achieves an 81% success rate, outperforming waveform, ASR, and compact audio-native baselines, and reaches 70% sound-causal success across four real-world Franka tasks. These results show that robust sound-centric manipulation requires not only native audio input, but also causal auditory persistence and explicit temporal grounding. Code and videos are available at https://hear.irmv.top .

Chang Nie, Tianchen Deng, Guangming Wang et al. · 0 citations
#machine learning Preprint Sep 2026

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support for each generated token using target-token-specific vocabulary projections of the encoded audio representations, measures frequency-band relevance through controlled spectral occlusion, and combines the two signals into a spectro-temporal relevance map. We evaluate STAG against ten post-hoc explanation methods across four grounding benchmarks, where it achieves the best event-localization performance on every dataset, and apply it to eight audio-language backbones without parameter updates. Counterfactual deletion further shows that removing the identified evidence selectively reduces confidence in the corresponding event and frequently removes it from the regenerated caption. These results provide behavioral support for the faithfulness and selectivity of the explanations.

Lucia Cascone, V. Fraenza, Michele Nappi et al. · 0 citations
Preprint Aug 2026

FATE: Frame-Level Audio-Visual Temporal Embedding

When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame-level Audio-visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame-level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross-video semantic and within-video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero-shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at \texttt{https://github.com/guankaisi/FATE}.

Kaisi Guan, Bingzi Zhang, Xihua Wang et al. · 0 citations
Preprint Aug 2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such problems as tool-interactive audio reasoning and develop SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools. To support this capability, we construct HIU-Corpus, comprising 65,492 interaction trajectories and 507.6 hours of audio across 24 tasks, 8 skills and 9 tools. SpeechAgent-R first learns structured interaction behaviors through trajectory-based supervised fine-tuning and then improves its decisions through multi-turn reinforcement learning. We further introduce HIU-Bench to jointly evaluate task performance, interaction quality and generalization to diverse task settings. It contains 1,395 samples across 56 tasks, including in-distribution (ID) and out-of-distribution (OOD) splits with substantial shifts in tool usage and workflow composition. SpeechAgent-R achieves 84.17 on ID tasks and 70.94 on OOD tasks, improving over the base model under the same agent harness by 15.40 and 14.23 points. These results demonstrate that learning skill and tool coordination improves audio agents'ability to handle diverse task settings and adaptive tool interactions.

Yuwen Wang, Tian-Hao Zhang, Ming Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.