
IS
Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
· 1 min read
ResearcharXiv cs.CL
What, When, and How: Audio Description as Constrained Global Optimization
arXiv:2609.30121v1 Announce Type: new
Abstract: Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.
Original source
This story was published by arXiv cs.CL and written by Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


