AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Some Math behind Neural Tangent Kernel

Collected Oct 1, 2026

Lilian Weng published a blog post in September 2022 titled "Some Math behind Neural Tangent Kernel" on Lil'Log, presenting a math-intensive deep dive into the neural tangent kernel (NTK) introduced by Jacot et al. in 2018.

The post focuses on a small number of core papers rather than a broad literature review. Weng notes that later works modifying or expanding the theory are not covered.

According to the post, neural networks are over-parameterized and can often fit data with near-zero training loss and decent generalization, with optimization consistently producing similarly good outcomes even when parameters exceed training data points. NTK is described as a kernel explaining the evolution of neural networks during training via gradient descent, offering insights into why sufficiently wide networks consistently converge to a global minimum when minimizing an empirical loss.

The post covers background material including vector-to-vector derivatives, differential equations, the Central Limit Theorem, Taylor expansion, kernel methods, and Gaussian processes, then proceeds to the NTK definition, infinite-width networks, the connection with Gaussian processes, deterministic NTK, linearized models, and lazy training.

It references Lee & Bahri et al. (2018) on deep neural networks as Gaussian processes, Chizat et al. (2019) on lazy training, Lee & Xiao et al. (2019) on wide networks evolving as linear models, and Arora et al. (2019) for a proof requiring only sufficiently large minimum width rather than all hidden layers being infinitely wide.

Weng states the goal is to present all the math behind NTK in a clear, easy-to-follow format and invites readers to report mistakes for correction.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Neural networks are well known to be over-parameterized and can often easily fit data with near-zero training loss with decent generalization performance on test dataset. Although all these parameters are initialized at random, the optimization process can consistently lead to similarly good outcomes. And this is true even when the number of model parameters exceeds the number of training data points. Neural tangent kernel (NTK) ( Jacot et al. 2018 ) is a kernel to explain the evolution of neural networks during tr