SyncAI.news, a Varaisys broadcasting
Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR
DM

Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Aditya Joshi, Akshay Agarwal, Jasabanta Patro

· 1 min read

ResearcharXiv cs.CL

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

arXiv:2605.29637v2 Announce Type: replace Abstract: Large language models often exhibit a substantial gap between their performance in English and in lower-resourced languages on equivalent knowledge queries---a cross-lingual consistency issue that remains underexplored for Indian languages and their code-mixed counterparts. To study this gap, we introduce IndicKLAR, an Indic extension of the KLAR-CLC benchmark covering 18 of the 22 scheduled Indian languages. For 11 widely used language pairs, we additionally provide code-mixed variants. Both monolingual and code-mixed inputs verified by native speakers. This three-way alignment enables us to examine how knowledge recall consistency varies across English, code-mixed, and native Indian language inputs. Across nine open-weight models, we find that the accuracy gap between native-language and English inputs can reach $\sim$0.50, while code-mixed inputs substantially reduce this gap, bringing performance within $\sim$0.05 of English without any model-level intervention. Motivated by this finding, we evaluate several prompting strategies that differ in how explicitly language conversion is exposed: a two-stage translate-then-answer setup, a one-stage joint translation-and-answer prompt, and Translate-in-Thought (TinT)---a single-step strategy in which the model internally converts the input and outputs only the final answer. Across the native $\rightarrow$ code-mixed $\rightarrow$ English performance trajectory, we observe a consistent flip point---the transition from incorrect to correct prediction---between the native and code-mixed settings. Notably, this pattern holds both when the code-mixed representation is explicitly provided as input or when the model is prompted to convert internally using TinT.

Original source

This story was published by arXiv cs.CL and written by Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Aditya Joshi, Akshay Agarwal, Jasabanta Patro. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News