SyncAI.news, a Varaisys broadcasting
Stable initialization without the CLT
SK

Simon Kuang, Kyle Chickering, Xinfan Lin

· 1 min read

ResearcharXiv cs.LG

Stable initialization without the CLT

arXiv:2609.30633v1 Announce Type: new Abstract: Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For networks with sine activations, we derive the uniform-phase initialization, which obviates distributional approximation and fully decouples the layers. Ours is the first work to use the sine function's periodic symmetry. Models trained with the uniform-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting. We find that our untuned models are competitive with the best-tuned baselines from previous work and support $\mu$P width scaling.

Original source

This story was published by arXiv cs.LG and written by Simon Kuang, Kyle Chickering, Xinfan Lin. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News