
YD
Yehoshua Dissen, Joseph Keshet, Eduard Golshtein
· 1 min read
ResearcharXiv cs.LG
Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization
arXiv:2609.39162v1 Announce Type: cross
Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.
Original source
This story was published by arXiv cs.LG and written by Yehoshua Dissen, Joseph Keshet, Eduard Golshtein. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


