Skip to content
Preprint

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

Aug 2026 · 0 citations · 78 references
Computer Science

TL;DR

This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.

Abstract

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the eval...

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al. · 2 citations
Preprint Aug 2026

Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled set...

Yizhou Zhang, Wangjin Zhou, Xin Gu et al. · 1 citation
Preprint Aug 2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.

Tony Alex, Wish Suharitdamrong, Sara Atito et al. · 0 citations
Preprint Aug 2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools, is developed and HIU-Bench is introduced to jointly evaluate task performance, interaction quality and generalization to diverse task settings.

Yuwen Wang, Tian-Hao Zhang, Ming Cai et al. · 0 citations
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

AudioICL-Bench is introduced, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge, and its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations
#machine learning Preprint Aug 2026

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

This work introduces MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs, and finds that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably.

Yi-Ze Li, Ning-Yuan Yang, Sile Yin et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.