SyncAI.news, a Varaisys broadcasting
Queer inclusion in speech datasets: An audit and taxonomy of practical tensions
BS

Brooklyn Sheppard, Anaelia Ovalle, Adina Williams, Levent Sagun

· 1 min read

ResearcharXiv cs.AI

Queer inclusion in speech datasets: An audit and taxonomy of practical tensions

arXiv:2609.25491v1 Announce Type: new Abstract: In this paper, we examine speech datasets for their inclusion of LGBTQIA+, or queer, voices and provide a taxonomy of tensions to better understand why there is a lack of such voices in current speech technology datasets. Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement. We take this community as a case study to consider what challenges and tensions are associated with collecting speech data from marginalized communities. For comparison, we audit an additional two datasets from the speech sciences that were created by, for, and with the queer community. We note that many customs in speech dataset collection efforts in AI and speech technology research may conflict with values emphasized in participatory approaches with marginalized communities, and provide a taxonomy describing these tensions.

Original source

This story was published by arXiv cs.AI and written by Brooklyn Sheppard, Anaelia Ovalle, Adina Williams, Levent Sagun. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News