This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.
Abstract
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of segmenting continuous audio streams into thematically coherent chapters, as a demanding and commercially consequential setting that exposes this gap. Chapterization is challenging because boundaries are defined less by objective acoustic events than by subjective editorial judgment, requiring models to reason sequentially over long acoustic contexts and approximate creator-authored boundary decisions. We present AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning. To support training and evaluation, we curate three datasets: AudioChaps-Alignment, derived from creator-annotated chapter boundaries on YouTube; AudioChaps-CoT, which provides structured supervision for well-formatted, high-quality, and evidence-grounded boundary reasoning; and AudioChaps-Eval, a held-out benchmark for audio chapterization. Applying GRPO directly without a Supervised Fine-Tuning (SFT) cold start, AudioChaps-R1-Zero already improves average F1 by 33 points over the state-of-the-art LALM Audio-Flamingo-3-Think. The AudioChaps framework produces our final aligned LALM, AudioChaps-R1, which improves average F1 by 49 points. These results demonstrate that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media. Our code, models, and dataset resources will be released upon acceptance at https://github.com/ta012/AudioChaps.
This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.
Wen-Jun Huang, Q. Chu, Tiger Shao et al.· 0 citations
This work proposes an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles, and shows that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and aco...
Liang-Yuan Wu, Sripathi Sridhar, M. Cartwright et al.· arXiv.org· 0 citations
SonicCaps is introduced, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text and training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves au...
Zineb Lahrichi, Marc Ferras, G. Richard et al.· 0 citations
This paper describes the system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline and obtains 90.92% accuracy on the final official evaluation set.
Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist...
Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra et al.· 0 citations
This work studies how to make LLMs natively generate TTS-friendly text, which is frame as a preference alignment problem: instead of relying on downstream rewriting modules, this work directly align LLMs to generate text optimized for spoken delivery.
Thibaut Thonet, Jos Rozen, Laurent Besacier· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.