
YE
Yasas Ekanayaka, Deshan Sumanathilaka
· 1 min read
ResearcharXiv cs.CL
A Character-Level Neural Approach to Sinhala Sandhi Splitting
arXiv:2609.36131v1 Announce Type: new
Abstract: Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-match accuracy (82.08\% character-level accuracy), well below the 94.00\% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.
Original source
This story was published by arXiv cs.CL and written by Yasas Ekanayaka, Deshan Sumanathilaka. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


