SyncAI.news, a Varaisys broadcasting
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
ZS

Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Hui Shen, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo

· 1 min read

ResearcharXiv cs.CL

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

arXiv:2608.12149v3 Announce Type: replace Abstract: We present the first systematic study of massive activations (MAs) in layer-interleaved Hybrid linear attention large language models (HLA LLMs), examining their architectural organization, training-time emergence, underlying mechanisms, and functional significance. Across five linear attention architectures, six hybridization configurations, and five input domains, we identify two architecture-aligned morphologies: pre-attention spikes (PAS) immediately before full attention and inter-spike plateaus (ISP) persisting through intervening linear attention layers. Denser full attention increasingly connects PAS through ISP, approaching the persistent MAs of conventional Transformers. This organization also recurs across 12 public checkpoints spanning 1.2B-397B parameters, covering linear attention and state-space hybrids. Controlled pretraining of Gated DeltaNet (GDN) hybrids up to 1.3B reveals early emergence and consolidation of both morphologies, alongside asymmetric gating effects. Specifically, full attention output gates strongly attenuate MA magnitudes without eliminating their organization, whereas removing GDN output gates yields modest amplification. Mechanistically, we develop a shared systematic-outlier account: PAS follows a localized write-sink-cancel process, while ISP is consistent with delayed cancellation. Functionally, our interventions show that deleting only the four largest-magnitude PAS coordinates at each full attention input reduces mean downstream accuracy by 21.9%-63.6% relative to normal inference. Moreover, reference-conditioned spike-to-plateau connection consistently improves mean real-world retrieval accuracy, yielding relative gains of 1.1%-12.6% without retraining. Our code is available at https://github.com/StartLuxLabs/Massive-Activations-HLA.

Original source

This story was published by arXiv cs.CL and written by Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Hui Shen, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News