SyncAI.news, a Varaisys broadcasting
A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis
TR

Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo

· 1 min read

ResearcharXiv cs.CL

A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis

arXiv:2609.30075v1 Announce Type: new Abstract: Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's $\rho\!=\!-0.53$) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset ($\rho\!=\!-0.34$). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.

Original source

This story was published by arXiv cs.CL and written by Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News