Scaling Laws, Carefully
Lilian Weng published a technical overview of scaling laws in deep learning, describing them as a framework for the relationship between compute, loss, model size and data, and, at their core, about allocating compute optimally between model size (N) and dataset size (D). The piece opens with early work on loss predictability: Amari et al. (1992) derived four learning curves via a Bayesian approach, and Hestness et al. (2017) observed across neural machine translation, image classification, language modeling and speech recognition that generalization error scales as a power law, with architecture changing the offset but not the exponent. Rosenfeld et al. (2020) modeled error jointly in N and D across architectures including ResNet, WRN, LSTM and Transformer.
The article then covers Kaplan et al. (2020), which popularized scaling laws for language models. Using models from 768M to 1.5B non-embedding parameters and 22M to 23B tokens, Kaplan et al. found loss scales as a power law with N, D and compute C individually, that larger models are more sample-efficient, and that given fixed compute it is more efficient to train a very large model and stop before convergence. They reported N_opt proportional to C^0.73, suggesting a 10x compute increase scale model size ~5.5x and tokens only ~1.8x.
The Chinchilla paper (Hoffmann et al. 2022) disagreed. Scanning over 400 models from 70M to over 16B parameters and 5B to 500B tokens, and using three methods — fixing model sizes while varying token budgets, IsoFLOP profiles, and a parametric fit — it arrived at N_opt proportional to C^0.5, implying model size and tokens should scale at equal rates. As a demonstration, Chinchilla (70B parameters, 1.4T tokens) was trained under the same compute budget as Gopher (280B parameters, 300B tokens) and outperformed Gopher across the board, supporting the claim that many large models at the time were undertrained. Weng also covers scaling laws in the data-limited region, the question of why the relationship is a power law, and difficulties fitting scaling laws in practice, including a toy simulation.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between $N$ and $D$.