
XL
Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu
· 1 min read
ResearcharXiv cs.CL
BAIT: Boundary-Guided Disclosure Escalation LLM Jailbreaking via Self-Conditioned Reasoning
arXiv:2605.27110v2 Announce Type: replace-cross
Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that elicits malicious information through internal disclosure by target large language models (LLMs), instead of external feedback from judge LLMs. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge requests; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering. These findings underscore the necessity of safeguarding against internally driven reasoning trajectories that progressively cross safety boundaries.
Original source
This story was published by arXiv cs.CL and written by Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


