Skip to content
Preprint

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

Aug 2026 · 0 citations · 21 references
Computer Science Engineering

TL;DR

This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.

Abstract

Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of segmenting continuous audio streams into thematically coherent chapters, as a demanding and commercially consequential setting that exposes this gap. Chapterization is challenging because boundaries are defined less by objective acoustic events than by subjective editorial judgment, requiring models to reason sequentially over long acoustic contexts and approximate creator-authored boundary decisions. We present AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning. To support training and evaluation, we curate three datasets: AudioChaps-Alignment, derived from creator-annotated chapter boundaries on YouTube; AudioChaps-CoT, which provides structured supervision for well-formatted, high-quality, and evidence-grounded boundary reasoning; and AudioChaps-Eval, a held-out benchmark for audio chapterization. Applying GRPO directly without a Supervised Fine-Tuning (SFT) cold start, AudioChaps-R1-Zero already improves average F1 by 33 points over the state-of-the-art LALM Audio-Flamingo-3-Think. The AudioChaps framework produces our final aligned LALM, AudioChaps-R1, which improves average F1 by 49 points. These results demonstrate that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media. Our code, models, and dataset resources will be released upon acceptance at https://github.com/ta012/AudioChaps.

View source

Similar papers

Preprint Aug 2026

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.

Wen-Jun Huang, Q. Chu, Tiger Shao et al. · 0 citations

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

This work proposes an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles, and shows that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and aco...

Liang-Yuan Wu, Sripathi Sridhar, M. Cartwright et al. · 0 citations
#natural language process... Preprint Sep 2026

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is introduced, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text and training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves au...

Zineb Lahrichi, Marc Ferras, G. Richard et al. · 0 citations
#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes the system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge, which adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline and obtains 90.92% accuracy on the final official evaluation set.

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
Review Aug 2026

AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP

Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist...

Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra et al. · 0 citations
#natural language process... Preprint Sep 2026

Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation

This work studies how to make LLMs natively generate TTS-friendly text, which is frame as a preference alignment problem: instead of relying on downstream rewriting modules, this work directly align LLMs to generate text optimized for spoken delivery.

Thibaut Thonet, Jos Rozen, Laurent Besacier · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.