
YZ
Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu
· 1 min read
ResearcharXiv cs.CL
Deep Delta Learning
arXiv:2601.00417v5 Announce Type: replace-cross
Abstract: Transformer residual streams are updated by addition. A sufficiently expressive residual block can represent content replacement, but the residual update itself has no operation that reads, compares, and replaces content. We introduce Deep Delta Learning (DDL), which applies the delta rule over network depth. Each layer reads the residual state along a learned direction, compares the readout with a learned target, and writes a gated rank-1 correction back along the same direction. A closed gate gives the identity map, and a unit gate overwrites the selected readout with the target. DDL works with the usual vector state or with an expanded state that stores several value channels, while attention and MLP blocks keep the original model width. We pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset. At both scales, every DDL variant has lower validation loss and higher average one-shot accuracy than the additive baseline, and the best expanded variants raise that average by 0.91 and 1.18 points.
Original source
This story was published by arXiv cs.CL and written by Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


