
DF
Demi\'an Fraiman
· 1 min read
ResearcharXiv cs.LG
On the Expressive Power of Transformers for Contextual Relations
arXiv:2603.25860v4 Announce Type: replace-cross
Abstract: Transformers have revolutionized machine learning by making attention a central mechanism for modeling interactions within a context. Despite the central role of attention, the theoretical capabilities of Transformers for representing contextual relations remain unclear. In this work, we address this question by developing a mathematical framework based on probability and optimal transport. We view a text as a distribution of its representations and attention as a probabilistic relation between them. This perspective reveals a connection between attention normalization and optimal transport: standard softmax normalization produces conditional relations, while Sinkhorn normalization produces joint relations with prescribed marginals. Thus, both mechanisms provide structured probabilistic relations from attention scores. Under mild conditions, we establish universal approximation results for both settings. We show that Transformer architectures with Sinkhorn normalization can approximate arbitrary contextual relations represented as joint probabilities, while standard softmax Transformers can approximate arbitrary contextual relations represented as conditional probabilities. These results provide a mathematical characterization of the expressive power of Transformers for contextual relations and show how the choice of normalization determines the probabilistic structure of the relations represented by attention.
Original source
This story was published by arXiv cs.LG and written by Demi\'an Fraiman. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


