
KN
Khang Nhat Hoang Vo, Anh Trac Duc Dinh, Tai Tien Ta, Tho Quan
· 1 min read
ResearcharXiv cs.CL
Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech
arXiv:2609.20223v1 Announce Type: new
Abstract: We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC--AUC of \(0.9953\), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.
Original source
This story was published by arXiv cs.CL and written by Khang Nhat Hoang Vo, Anh Trac Duc Dinh, Tai Tien Ta, Tho Quan. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


