SyncAI.news, a Varaisys broadcasting
An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations
LW

Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes

· 1 min read

ResearcharXiv cs.CL

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

arXiv:2607.21424v2 Announce Type: replace Abstract: Recent advances in automated audio captioning (AAC) are driving a shift from monolithic sentences toward structured formats that disentangle acoustic and semantic properties, such as timestamped captions for different sound events. Such representations can support faceted sound search for creators and richer access to auditory information for Deaf and Hard of Hearing people. Yet, it remains unclear how to meaningfully evaluate these hybrid, structured captions. We propose an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles. The framework combines large language model (LLM) judges for semantic fields with deterministic metrics for temporal and acoustic attributes. To validate these metrics, we introduce controlled perturbations that apply typed, graded changes to ground-truth annotations. Results show that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and acoustic corruptions, enabling more reliable evaluation of structured captions.

Original source

This story was published by arXiv cs.CL and written by Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News