This paper describes a neurosymbolic architecture for learning to assemble novel structures using evidence from embodied conversations and task demonstrations. We focus on scenarios where an agent encounters, after deployment, semantic constraints on structures--in other words, constraints as to which part types and fe...
Jonghyuk Park, A. Lascarides, S. Ramamoorthy· 0 citations
This hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations, which makes better decisions than prompted LLMs about what to describe and when to describe it.
Igor Sterner, Mirella Lapata, A. Lascarides et al.· 0 citations
Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being to...
Igor Sterner, Mirella Lapata, A. Lascarides et al.· 2 citations· ⚡1
Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most narratives: the descriptions not only need to convey the story in the movie, but they must als...
Igor Sterner, A. Lascarides, Frank Keller· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.