The Transformer Family Version 2.0
Lilian Weng has published version 2.0 of "The Transformer Family," a major refactoring and expansion of her 2020 post on the same topic, roughly three years after the original. According to the post, the update restructures the hierarchy of sections and improves many sections with more recent papers. Version 2.0 is described as a superset of the old version and about twice its length.
The article begins with a notation table defining terms used throughout, including model size or hidden state dimension d, number of attention heads h, input sequence segment length L, and total attention layers N, noting N does not consider MoE. It also lists matrices for queries, keys, values, per-head weights, and the output weight, along with the self-attention matrix A and scalar attention scores.
The post covers Transformer basics, describing the vanilla Transformer (Vaswani, et al., 2017) as an encoder-decoder architecture commonly used in NMT models, with later simplified versions performing well in language modeling tasks such as encoder-only BERT or decoder-only GPT. It reviews attention and self-attention, scaled dot-product attention, and multi-head self-attention, including the concatenation and linear transformation of per-head outputs.
The section list also includes encoder-decoder architecture, positional encoding variants (sinusoidal, learned, relative, and rotary position embedding from Su et al. 2021), longer context approaches, context memory such as Transformer-XL (Dai et al., 2019) and Compressive Transformer (Rae et al. 2019), non-differentiable external memory including kNN-LM (Khandelwal et al. 2020) and SPALM (Yogatama et al. 2021), distance-enhanced attention scores, recurrence, adaptive modeling, efficient attention, sparse attention patterns, low-rank attention, and transformers for reinforcement learning.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment of that 2020 post — restructure the hierarchy of sections and improve many sections with more recent papers. Version 2.0 is a superset of the old version, about twice the length. Notations Symbol Meaning $d$ The model size / hidden state dimension / positional encoding size. $h$ The number of heads in multi-head attentio