SyncAI.news, a Varaisys broadcasting
Benchmarking Automatic Speech Recognition Tools for Iberian Languages
FL

Fernando L\'opez, Pablo G\'omez, David Solans, Paulo Villegas, Jordi Luque

· 1 min read

ResearcharXiv cs.CL

Benchmarking Automatic Speech Recognition Tools for Iberian Languages

arXiv:2609.36920v1 Announce Type: new Abstract: Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.

Original source

This story was published by arXiv cs.CL and written by Fernando L\'opez, Pablo G\'omez, David Solans, Paulo Villegas, Jordi Luque. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News