This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.
Abstract
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.
This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the eval...
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al.· 2 citations
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled set...
Yizhou Zhang, Wangjin Zhou, Xin Gu et al.· 1 citation
This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.
Tony Alex, Wish Suharitdamrong, Sara Atito et al.· 0 citations
SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools, is developed and HIU-Bench is introduced to jointly evaluate task performance, interaction quality and generalization to diverse task settings.
Yuwen Wang, Tian-Hao Zhang, Ming Cai et al.· 0 citations
AudioICL-Bench is introduced, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge, and its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to...
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al.· 0 citations
This work introduces MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs, and finds that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably.
Yi-Ze Li, Ning-Yuan Yang, Sile Yin et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.