
AG
Ankur Garg, Xuemin Yu, Hassan Sajjad, Samira Ebrahimi Kahou
· 1 min read
ResearcharXiv cs.AI
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
arXiv:2506.20040v4 Announce Type: replace-cross
Abstract: Interpreting language models remains challenging due to the existence of residual stream, which linearly mixes and duplicates features across adjacent layers, causing single-layer analyses to miss this cross-layer structure. Cross-layer sparse autoencoders (SAEs) address layer mixing but operate in continuous space, where concepts split across many neurons without clear boundaries. We introduce Cross-Layer Vector Quantized-Variational Autoencoder (CLVQ-VAE), a novel framework which maps representations from a lower layer to a higher layer through a discrete vector-quantization bottleneck, collapsing duplicated residual-stream features into compact, interpretable concept vectors. Our approach combines top-k temperature-based sampling with exponential moving average (EMA) codebook updates, providing controlled exploration of the discrete latent space while maintaining codebook diversity. Across both encoder- and decoder-based models on ERASER-Movie, Jigsaw, and AGNews, CLVQ-VAE outperforms clustering, single-layer vector quantized-variational autoencoder (VQ-VAE), and sparse autoencoder (SAE) baselines across three evaluation axes: removing identified concepts drops downstream probe accuracy by up to 93%, LLM judges rank our concepts first in 66.7% of comparisons, and human annotators recover model predictions from our visualizations with 78% accuracy versus 54% for clustering.
Original source
This story was published by arXiv cs.AI and written by Ankur Garg, Xuemin Yu, Hassan Sajjad, Samira Ebrahimi Kahou. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


