Skip to content

What, When, and How: Audio Description as Constrained Global Optimization

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

This hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations, which makes better decisions than prompted LLMs about what to describe and when to describe it.

Abstract

Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.

View source

Similar papers

Preprint Aug 2026

REFRAMED: Towards Realistic Audio Description Generation for Movies

Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being to...

Igor Sterner, Mirella Lapata, A. Lascarides et al. · 2 citations · ⚡1
Preprint Sep 2026

From Visual Cues to Spoken Narration: Rethinking Audio Description

Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achieve the best user experience. Prior work has largel...

Akshita Gupta, Aditya Arora, Federico Tombari et al. · 1 citation
Preprint Aug 2026

Comparing British and American Audio Description of Movies

Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most narratives: the descriptions not only need to convey the story in the movie, but they must als...

Igor Sterner, A. Lascarides, Frank Keller · 1 citation
Preprint Aug 2026

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

This paper provides an evaluation and design framework for comparing representation-model pairs and shows that RVQ's residual order gives ordered capacity but not ordered semantics, and that AudioLM's semantic-versus-acoustic cascade is one explicit placement of this boundary rather than a universal template.

Dongchao Yang · 0 citations
Preprint Aug 2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...

Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al. · 3 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.