SyncAI.news, a Varaisys broadcasting
SteganoBackdoor: Evading Data-Poisoning Defenses via Steganographic Backdoors
EX

Eric Xue, Ruiyi Zhang, Pengtao Xie

· 1 min read

ResearcharXiv cs.CL

SteganoBackdoor: Evading Data-Poisoning Defenses via Steganographic Backdoors

arXiv:2511.14301v4 Announce Type: replace-cross Abstract: Transformer-based models are highly susceptible to backdoor attacks via supervised fine-tuning (SFT). To red-team existing data-poisoning defenses, prior work has increasingly focused on stylized triggers, synthetic artifacts, and token-level perturbations designed to evade detection. However, this trend has shifted threat models away from naturally occurring semantic triggers and realistic low-budget poisoning settings. Addressing this gap, we introduce SteganoBackdoor, an optimization-based framework that transforms semantic-trigger seeds through autoregressive token replacement, sequentially minimizing embedding overlap with the inference-time trigger while preserving a strong per-sample training-time payload. The resulting SteganoPoisons maintain linguistic fluency and encode the payload across ordinary tokens, such that no individual token carries a concentrated signal and the full payload instead emerges from their exact combination and ordering. Across 18 encoder-based and decoder-only models spanning 120M to 14B parameters, SteganoBackdoor achieves high attack success under sub-percent poisoning budgets and exposes limitations in existing data-poisoning defenses.

Original source

This story was published by arXiv cs.CL and written by Eric Xue, Ruiyi Zhang, Pengtao Xie. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News