SyncAI.news, a Varaisys broadcasting
A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training
IZ

Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky

· 1 min read

ResearcharXiv cs.AI

A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

arXiv:2605.14605v3 Announce Type: replace-cross Abstract: Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typically evaluated against attacks with a fixed training budget, even though an attacker who holds the weights can simply train for longer. We ask whether current defenses withstand this simplest escalation. Surveying fifteen recent defenses, we find that they share a common weakness: each is built around a limited model of the attacker, such as a bounded perturbation, a short simulated attack, or a trained link between harmful and benign behavior, and nothing enforces that protection once the weights are released. We then test six representative defenses on four open-weight models by continuing the same harmful-only fine-tuning attack for three epochs and measuring harmfulness and capability along the way. In all 72 defended runs, the model is more harmful at the end of training than at release, and on Llama-3.1 at the highest learning rate the defended models end almost as harmful as the undefended one. The defenses are not equally weak: one defense kept harmfulness low on one model, and some attacks recovered harmfulness only at the cost of general capability. Current defenses can delay or disrupt malicious fine-tuning, but in most cases their measured resistance does not persist under continued training, and they should not yet be treated as durable protection.

Original source

This story was published by arXiv cs.AI and written by Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News