Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most narratives: the descriptions not only need to convey the story in the movie, but they must also fit into gaps between dialogue and they need to conform to guidelines that exist in each region. In this work, we compare audio description created in the United Kingdom against audio description created in the United States. We use guidelines written for these two regions, alongside the impressions from a practitioner in the field, to motivate specific hypotheses about the differences. We test these hypotheses against our pre-existing corpus, which provides both human-authored American and British audio description for each of 206 movies. Results provide quantitative evidence to uphold all tested hypotheses, including differences in lexicon, the use of the progressive aspect, the use of passive constructions, the use of subjective adjectives and modifiers, when characters are named, how scenes are cued, and degree of overlap with movie dialogue and music. Our work offers a quantitative lens into the narrative technique of audio description.
Introduction: Experience as Contested Territory
When Gardening Australia became the first program to broadcast with audio description (AD) on the ABC in 2020, the national broadcaster received complaints from sighted viewers about why their televisions were suddenly talking to them. This accessibility feature describes...
Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achieve the best user experience. Prior work has largel...
Akshita Gupta, Aditya Arora, Federico Tombari et al.· 1 citation
Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being to...
Igor Sterner, Mirella Lapata, A. Lascarides et al.· 2 citations· ⚡1
Over the course of the twentieth century, sound and the moving image gradually came together to form productions combining the two media, such as cinema and music video art. However, few studies analyse their interaction from the point of view of the audio-spectator, and none look at their impact on the development of...
J. Moreau, Éric Tortochot, P. Terrien et al.· Travaux interdisciplinaires...· 0 citations
This hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations, which makes better decisions than prompted LLMs about what to describe and when to describe it.
Igor Sterner, Mirella Lapata, A. Lascarides et al.· 0 citations