SyncAI.news, a Varaisys broadcasting
mu-bench: A Multilingual Utterance Transcription Benchmark
AL

Andrea Li (UC Berkeley), Soham Ray (Sierra AI)

· 1 min read

ResearcharXiv cs.CL

mu-bench: A Multilingual Utterance Transcription Benchmark

arXiv:2609.32082v1 Announce Type: new Abstract: Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an AI banking agent in English, Spanish, Turkish, Vietnamese, and Mandarin, centered on form-field inputs such as names, email addresses, and confirmation codes. We release Utterance Error Rate (UER), an LLM judge of whether a transcript preserves meaning, calibrated against human raters, together with an LLM normalizer that makes WER comparable across providers' output formats. On 1,847 human-rated transcripts, UER agrees with annotators at $\kappa$ = 0.78, versus 0.53 for exact-match WER on normalized text. We rank six commercial providers on a public leaderboard; the best reaches 11.9% UER, and Mandarin is hardest for all six.

Original source

This story was published by arXiv cs.CL and written by Andrea Li (UC Berkeley), Soham Ray (Sierra AI). SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News