Separation of timescales controls feature learning and overfitting in large neural networks
Pierfrancesco Urbani (IPhT, Gif-sur-Yvette)
Understanding how large overparameterized networks can extract relevant features from meaningful data while simultaneously be so expressive to perfectly interpolate pure noise is a conceptual problem central in machine learning. To address this, we analyze the out-of-equilibrium, non-linear, high-dimensional training dynamics of overparametrized two-layer neural networks using dynamical mean field theory.
Our findings reveal a separation of timescales in training, leading to several key observations:
– A slow timescale linked to the growth in model complexity;
– An inductive bias towards low complexity when the initial model complexity is sufficiently small;
– A dynamical decoupling between feature learning and overfitting phases;
– A non-monotonic trend in test error, characterized by a « feature unlearning » regime at later stages of training.
Joint work with Andrea Montanari.
