
DL
Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
· 1 min read
ResearcharXiv cs.CL
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
arXiv:2609.03322v2 Announce Type: replace
Abstract: Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints, analyze layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores show the strongest pooled associations with activation-patching recovery under token substitution and shuffling, although these associations do not isolate copying-specific effects. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.
Original source
This story was published by arXiv cs.CL and written by Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


