
SK
Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh
· 1 min read
ResearcharXiv cs.CL
A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition
arXiv:2609.30160v1 Announce Type: new
Abstract: Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation of the wildcard weights with a predictive-entropy-indexed schedule, which performs comparably while reducing dependence on training length. Combining this schedule with the hybrid graph gives the lowest mean word error rate (WER) on every corpus and a 9.45% average relative WER reduction over CTC. Independent validator transcriptions show that token-level models place significantly more wildcard-bypass probability than CTC on disputed characters, indicating that token-level tolerance targets localized transcript ambiguity.
Original source
This story was published by arXiv cs.CL and written by Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


