AI Insight
This theoretical study analyzes scaling laws in shallow neural networks operating in the feature learning regime, specifically examining quadratic and diagonal network architectures. By connecting these networks to matrix compressed sensing and LASSO methods, the researchers derive a phase diagram showing how prediction error scales with dataset size and regularization strength, identifying distinct scaling regimes and transitions between them. The work establishes a rigorous theoretical link between the power-law distribution of network weight spectra and generalization performance, providing mathematical foundations for empirical observations in deep learning.
Why it matters
Understanding scaling laws theoretically helps predict how neural networks will perform as they grow larger or receive more training data, potentially improving resource allocation and architecture design decisions. The connection between weight spectrum properties and generalization could lead to better diagnostic tools for assessing model quality and new regularization strategies.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveraging connections with matrix compressed sensing and LASSO, we derive a detailed phase diagram for the scaling exponents of the excess risk as a function of sample complexity and weight decay. This analysis uncovers crossovers between distinct scaling regimes and plateau behaviors, mirroring phenomena widely reported in the empirical neural scaling literature. Furthermore, we establish a precise link between these regimes and the spectral properties of the trained network weights, which we characterize in detail. As a consequence, we provide a theoretical validation of recent empirical observations connecting the emergence of power-law tails in the weight spectrum with network generalization performance, yielding an interpretation from first principles.
Source: Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime