SyncAI.news, a Varaisys broadcasting
Gaussian Equivalence for Multi-Head Self-Attention
TH

Tomohiro Hayase, Ryo Karakida

· 1 min read

ResearcharXiv cs.LG

Gaussian Equivalence for Multi-Head Self-Attention

arXiv:2610.10033v1 Announce Type: cross Abstract: A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.

Original source

This story was published by arXiv cs.LG and written by Tomohiro Hayase, Ryo Karakida. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News