
MB
Marco Biroli
· 1 min read
ResearcharXiv cs.AI
Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
arXiv:2609.32116v1 Announce Type: new
Abstract: Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independent augmentations of an unsafe prompt and sampling $M$ completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in $N$, which we challenge. The exponent drifts with $N$, with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget ($N, M$) attack surface as well as its dependence on the generation temperature $T$. We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire ($N, M$) attack surface. They extrapolate predictions from $N \leq 100$ to $N = 10^4$, collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.
Original source
This story was published by arXiv cs.AI and written by Marco Biroli. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


