
YZ
Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Bo Li
· 1 min read
ResearcharXiv cs.CL
BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs
arXiv:2605.00689v2 Announce Type: replace
Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety across diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benchmarks largely rely on general risk taxonomies and machine-translated data, limiting evaluation to predefined risk categories and providing insufficient coverage of region-specific regulatory requirements and cultural contexts. To bridge these gaps, we introduce BabelSafe, a policy-grounded multilingual safety benchmark covering 13 language settings. BabelSafe is constructed from regional regulatory sources, with risk categories and fine-grained rules extracted from jurisdiction-specific regulatory documents directly used to guide the generation of multilingual safety data. During data generation, we further incorporate region-specific cultural contexts, enabling regulation-grounded and culturally contextualized evaluation across languages. Building on BabelSafe, we develop BabelGuard, a Diffusion Large Language Model (dLLM)-based guardrail model that supports multilingual safety judgment and policy-conditioned safety assessment. BabelGuard has two variants, a lightweight 1.5B model for fast `safe/unsafe' classification and a more capable 7B model for customizable policy-conditioned safety checking with detailed explanations. We evaluate BabelGuard against 11 strong guardrail baselines on 6 existing multilingual safety benchmarks and BabelSafe, demonstrating the strong performance of BabelGuard across these evaluation settings. We hope that BabelSafe and BabelGuard can help advance the development of regulation-aware and culturally contextualized multilingual guardrail systems.
Original source
This story was published by arXiv cs.CL and written by Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Bo Li. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


