扶摇AI知识笔记AI 前沿知识库
首页 / 产业观察 / 正文
产业观察

Double descent is the principle of least action

来源:arXiv cs.AI 论文速递 约 2083 字
arXiv cs.AI
转载

本文转载自 arXiv cs.AI,版权归原作者及原发布平台所有。本站仅作知识整理与转载分享,如涉版权问题请联系客服删除。

01核心要点

  • The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon.
  • We explain the phenomenon with statistical mechanics.
  • The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzman

02正文全文

Abstract:The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the $d$ degrees of freedom in shares of $T/2$, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the $L^2$ norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing $d$, effectively increasing weight regularization.

Ancillary files (details):

lean/DoubleDescent/Assumptions.lean

lean/DoubleDescent/Axioms.lean

lean/DoubleDescent/Basic.lean

lean/DoubleDescent/CompleteSquare.lean

lean/DoubleDescent/FiniteTime.lean

lean/DoubleDescent/NormDrop.lean

lean/DoubleDescent/Regimes.lean

scripts/fig_graphical_abstract.py

scripts/fig_polynomial.py

(12 additional files not shown)

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

03原文直达

本文内容转载自 arXiv cs.AI,如需查看原排版、配图与最新修订,请访问原始出处。

阅读原文(arXiv cs.AI)

下载论文 PDF

正在校验阅读权限…
RELATED

相关阅读

更多 产业观察