Explorerโ€บArtificial Intelligenceโ€บAI
Research PaperResearchia:202609.17057

Double descent is the principle of least action

Congzhou M Sha

Abstract

The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equal...

Submitted: September 17, 2026Subjects: AI; Artificial Intelligence

Description / Details

The test error of a model plotted against its number of parameters dd falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature TT, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the dd degrees of freedom in shares of T/2T/2, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the L2L^2 norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing dd, effectively increasing weight regularization.


Source: arXiv:2609.19076v1 - http://arxiv.org/abs/2609.19076v1 PDF: https://arxiv.org/pdf/2609.19076v1 Original Link: http://arxiv.org/abs/2609.19076v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 17, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark