
JL
Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan
· 1 min read
ResearcharXiv cs.CL
Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs
arXiv:2609.22144v1 Announce Type: new
Abstract: Preserving safety alignment during large language models fine-tuning is critical, however, recent studies have demonstrated that even benign fine-tuning data may contain safety-degrading samples that silently undermine safety alignment. Existing approaches typically identify such samples using representations from a single safety-sensitive layer. While this assumption has shown effectiveness in monolingual settings, its validity for multilingual models remains unclear due to potential cross-lingual differences in representation patterns. Through a cross-lingual analysis, we show that sensitive layers are only partially shared across languages, with safety-relevant signals often distributed across multiple layers. Motivated by these observations, we propose MMSAFE, a multi-layer framework for multilingual safety-degrading data identification that captures both shared and language-specific safety signals. Extensive experiments across multiple models, languages, and safety benchmarks demonstrate that MMSAFE reduces the average harmful-response ratio by 60% compared with random filtering and achieves stronger average performance than the strongest single-layer baseline, demonstrating the effectiveness of multi-layer modeling for robust multilingual safety alignment.
Original source
This story was published by arXiv cs.CL and written by Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


