
LW
Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes
· 1 min read
ResearcharXiv cs.CL
An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations
arXiv:2607.21424v2 Announce Type: replace
Abstract: Recent advances in automated audio captioning (AAC) are driving a shift from monolithic sentences toward structured formats that disentangle acoustic and semantic properties, such as timestamped captions for different sound events. Such representations can support faceted sound search for creators and richer access to auditory information for Deaf and Hard of Hearing people. Yet, it remains unclear how to meaningfully evaluate these hybrid, structured captions. We propose an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles. The framework combines large language model (LLM) judges for semantic fields with deterministic metrics for temporal and acoustic attributes. To validate these metrics, we introduce controlled perturbations that apply typed, graded changes to ground-truth annotations. Results show that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and acoustic corruptions, enabling more reliable evaluation of structured captions.
Original source
This story was published by arXiv cs.CL and written by Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


