Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

작성자

카테고리:

← 피드로
arXiv cs.AI · Leonardo Defilippis, Yizhou Xu, Julius Girardin, Emanuele Troiani, Vittorio Erba, Lenka Zdeborov'a, Bruno Loureiro, Florent Krzakala · 2026-08-07 AI

[Submitted on 29 Sep 2025 (v1), last revised 6 Aug 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveraging connections with matrix compressed sensing and LASSO, we derive a detailed phase diagram for the scaling exponents of the excess risk as a function of sample complexity and weight decay. This analysis uncovers crossovers between distinct scaling regimes and plateau behaviors, mirroring phenomena widely reported in the empirical neural scaling literature. Furthermore, we establish a precise link between these regimes and the spectral properties of the trained network weights, which we characterize in detail. As a consequence, we provide a theoretical validation of recent empirical observations connecting the emergence of power-law tails in the weight spectrum with network generalization performance, yielding an interpretation from first principles.

Submission history

From: Yizhou Xu [view email]
[v1] Mon, 29 Sep 2025 14:58:13 UTC (1,511 KB)
[v2] Thu, 4 Jun 2026 17:25:40 UTC (1,626 KB)
[v3] Thu, 6 Aug 2026 09:40:55 UTC (1,624 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2509.24882

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다