SyncAI.news, a Varaisys broadcasting
Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA
WH

Wenxuan He, Yunpeng Li, Zewei Li, Yongke Yang, Yuze Li, Yin Cao, Shan Liang

· 1 min read

ResearcharXiv cs.LG

Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

arXiv:2608.19084v2 Announce Type: replace Abstract: Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.

Original source

This story was published by arXiv cs.LG and written by Wenxuan He, Yunpeng Li, Zewei Li, Yongke Yang, Yuze Li, Yin Cao, Shan Liang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News