AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Are Deep Neural Networks Dramatically Overfitted?

Collected Oct 1, 2026

Lilian Weng published an article asking whether deep neural networks are dramatically overfitted, motivated by the observation that a typical deep network has many parameters and can easily achieve perfect training error while still generalizing to out-of-sample data. The post was updated on 2019-05-27 to add a section on the Lottery Ticket Hypothesis.

The article surveys classic results on compression and model selection, including Occam's Razor, the Minimum Description Length principle, Kolmogorov Complexity, and Solomonoff's Inference Theory. It then covers the expressive power of deep learning models, the Universal Approximation Theorem, and a proof on the finite sample expressivity of two-layer neural networks.

Citing Zhang et al. (2017), the article states that a two-layer neural network with ReLU activations and 2n + d weights can represent any function on a sample of size n in d dimensions, and that deep networks can learn unstructured random noise perfectly, with results unchanged when regularization terms are added. It also notes that explicit regularization such as data augmentation, weight decay, and dropout is described in that work as neither necessary nor sufficient for reducing generalization error.

On the modern risk curve, the article cites Belkin et al. (2018) on a double-U-shaped bias-variance risk curve for deep neural networks, and mentions two proposed reasons: that parameter count is not a good measure of inductive bias, and that larger models may find interpolating functions with smaller norm. Weng writes that reproducing the smooth curve required careful attention to experimental details. The article also discusses intrinsic dimension (Li et al., 2018) and other complexity measures, such as degrees of freedom (Gao & Jojic, 2016) and prequential code (Blier & Ollivier, 2018), plus heterogeneous layer robustness.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

[Updated on 2019-05-27: add the section on Lottery Ticket Hypothesis.] If you are like me, entering into the field of deep learning with experience in traditional machine learning, you may often ponder over this question: Since a typical deep neural network has so many parameters and training error can easily be perfect, it should surely suffer from substantial overfitting. How could it be ever generalized to out-of-sample data points?