SyncAI.news, a Varaisys broadcasting
Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs
AI

Aaron Isidore Grace, Weiran Wang

· 1 min read

ResearcharXiv cs.AI

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

arXiv:2609.28879v1 Announce Type: cross Abstract: Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.

Original source

This story was published by arXiv cs.AI and written by Aaron Isidore Grace, Weiran Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News