SyncAI.news, a Varaisys broadcasting
Attention Kernels for Learning Maps Between Heavy-Tailed Measures
KH

Kailen Hargenrader, Edoardo Calvello, Bohan Chen

· 1 min read

ResearcharXiv cs.LG

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

arXiv:2610.00564v1 Announce Type: new Abstract: Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.

Original source

This story was published by arXiv cs.LG and written by Kailen Hargenrader, Edoardo Calvello, Bohan Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News