SyncAI.news, a Varaisys broadcasting
Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
NO

Navam Obeysekara, Nevidu Jayatilleke

· 1 min read

ResearcharXiv cs.CL

Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation

arXiv:2610.11389v1 Announce Type: new Abstract: Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source. This vulnerability is acute in low-resource settings like Sinhala-to-English, where weak cross-lingual alignment leads to hallucinations. This paper introduces a framework for reference-free hallucination detection in this language pair. We present a 45,000-sample synthetic dataset generated through a probabilistic chain of five linguistically motivated corruption strategies, with a semantic rescue mechanism that uses character-level similarity to distinguish hallucinations from morphological variants. We fine-tune mDeBERTa-v3 for token-level sequence labelling, reaching a token-level F1 of 0.841 +/- 0.001 over three seeds on a source-disjoint test set, and study a three-signal ensemble integrating neural risk scores, sequence log-probabilities, and cross-lingual semantic embeddings (LaBSE). A source-ablation control shows that the detector relies on the Sinhala source rather than on surface artefacts of the corruption process: shuffling or removing the source reduces sentence-level AUROC from 0.970 to chance. We benchmark eight NMT systems spanning five model families and find that detector firings vary by an order of magnitude across architectures.

Original source

This story was published by arXiv cs.CL and written by Navam Obeysekara, Nevidu Jayatilleke. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News