AI & Computational Science

How Shallow Neural Networks Learn Features at Different Scales

How the science connects

Neural networkCompressed sensingFeature learning

AI Insight

This theoretical study analyzes scaling laws in shallow neural networks operating in the feature learning regime, specifically examining quadratic and diagonal network architectures. By connecting these networks to matrix compressed sensing and LASSO methods, the researchers derive a phase diagram showing how prediction error scales with dataset size and regularization strength, identifying distinct scaling regimes and transitions between them. The work establishes a rigorous theoretical link between the power-law distribution of network weight spectra and generalization performance, providing mathematical foundations for empirical observations in deep learning.


Understanding scaling laws theoretically helps predict how neural networks will perform as they grow larger or receive more training data, potentially improving resource allocation and architecture design decisions. The connection between weight spectrum properties and generalization could lead to better diagnostic tools for assessing model quality and new regularization strategies.


Understand the Science

Neural network 59 articles Explore Concept → Compressed sensing Concept coming soon Feature learning Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveraging connections with matrix compressed sensing and LASSO, we derive a detailed phase diagram for the scaling exponents of the excess risk as a function of sample complexity and weight decay. This analysis uncovers crossovers between distinct scaling regimes and plateau behaviors, mirroring phenomena widely reported in the empirical neural scaling literature. Furthermore, we establish a precise link between these regimes and the spectral properties of the trained network weights, which we characterize in detail. As a consequence, we provide a theoretical validation of recent empirical observations connecting the emergence of power-law tails in the weight spectrum with network generalization performance, yielding an interpretation from first principles.

Source: Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime