SyncAI.news, a Varaisys broadcasting
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
EA

Enes Arda, Semih Cayci, Atilla Eryilmaz

· 1 min read

ResearcharXiv cs.LG

A Comparative Analysis of Attention versus State-Space Models for In-Context Learning

arXiv:2609.32341v1 Announce Type: new Abstract: Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.

Original source

This story was published by arXiv cs.LG and written by Enes Arda, Semih Cayci, Atilla Eryilmaz. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News