
GL
Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen
· 1 min read
ResearcharXiv cs.CL
Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
arXiv:2609.33634v1 Announce Type: new
Abstract: Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.
Original source
This story was published by arXiv cs.CL and written by Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


