
DK
David Kletz, Sandra Mitrovi\'c, Ljiljana Dolami\'c, Fabio Rinaldi
· 1 min read
ResearcharXiv cs.CL
Are Language Models Script-Aware?
arXiv:2610.08037v1 Announce Type: new
Abstract: Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
Original source
This story was published by arXiv cs.CL and written by David Kletz, Sandra Mitrovi\'c, Ljiljana Dolami\'c, Fabio Rinaldi. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


