SyncAI.news, a Varaisys broadcasting
FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs
SS

Shivam Saini, Eric Bezzam, Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen

· 1 min read

ResearcharXiv cs.AI

FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs

arXiv:2609.38897v1 Announce Type: cross Abstract: Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.

Original source

This story was published by arXiv cs.AI and written by Shivam Saini, Eric Bezzam, Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News