
GZ
Gongyue Zhang, Honghai Liu
· 1 min read
ResearcharXiv cs.LG
The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs
arXiv:2610.00403v1 Announce Type: new
Abstract: A network can fit its training examples while failing to recover the rule that generated their labels. We examine this separation in single-hidden-layer multilayer perceptrons (MLPs), using synthetic tasks that control interaction order and the presence of nuisance inputs. We establish elementary benchmark properties: pure parity contains no predictive lower-order marginals, admits an exact Bayes posterior, and can be represented on clean latent inputs by a width-$k$ ReLU network. Experiments then identify distinct optimization outcomes. In a matched order-2--4 sweep, SGD, Adam, and Muon all reach 100\% peak test accuracy at order two; at order three they reach 96.25\%, 50.87\%, and 76.82\%, respectively, while Muon reaches 99.21\% at order four. In a separate mixed-order task, freezing only the first-layer weights connected to independent nuisance inputs raises AdamW's epoch-10 accuracy from 44.73\% to 95.07\%. Removing the same inputs only at test time raises it to 48.38\%. Thus, nuisance-weight learning changes the training outcome beyond its immediate effect on prediction. Bias interventions expose a connection between target symmetry and shallow ReLU representations. In a compact signal-only regime, both SGD and Muon learn orders five through eight, with higher SGD peak accuracy at orders nine through eleven. Together, the results show how optimization and nuisance learning constrain the higher-order rules realized by a shallow network.
Original source
This story was published by arXiv cs.LG and written by Gongyue Zhang, Honghai Liu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


