
JK
Junhyeok Kim, Jinyeong Kim, Jae Wan Park, Seong Jae Hwang
· 1 min read
ResearcharXiv cs.CV
On the Necessity of Attention-FFN Split in Vision Transformers
arXiv:2610.10303v1 Announce Type: new
Abstract: The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.
Original source
This story was published by arXiv cs.CV and written by Junhyeok Kim, Jinyeong Kim, Jae Wan Park, Seong Jae Hwang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


