
SG
St\'ephane Galatolo, St\'ephane Chr\'etien
· 1 min read
ResearcharXiv cs.LG
Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
arXiv:2609.30274v1 Announce Type: cross
Abstract: Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most rudimentary, while still outperforming more recent algorithms at a number of learning tasks.
A major open question about the current methods used in deep learning is to understand their convergence properties. Following a line of previous works about the long time behavior of gradient-type algorithms, %and in particular the recent contributions from Azizian et al., we present a new approach for studying the asymptotic properties of a wide family of methods from an ergodic theoretical viewpoint. Our main results include a study of the expected time for a stochastic optimisation algorithm to reach a certain small neighborhood of a minimizer and show that this reaching time distributes exponentially around its average, which is given by the inverse of the stationary measure of the target. The assumptions on the Stochastic Gradient noise include the Gaussian and the Sub-Exponential assumptions.
Original source
This story was published by arXiv cs.LG and written by St\'ephane Galatolo, St\'ephane Chr\'etien. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


