AivexaNewsSearch
AI news for builders and product teamsChecked every hour

The Transformer Family

Collected Oct 1, 2026

Lilian Weng published a refactoring update to her post on the Transformer family on 2023-01-27, after almost three years, to incorporate a number of new Transformer models introduced since 2020. She directed readers to an enhanced version titled The Transformer Family Version 2.0 for that topic.

The original post covered notations; attention and self-attention; multi-head self-attention; the Transformer; Adaptive Computation Time; and techniques for improved attention span, including Transformer-XL's longer attention span, adaptive attention span, and localized attention span in the Image Transformer. It also covered approaches for lowering time and memory cost, including sparse attention matrix factorization in Sparse Transformers, locality-sensitive hashing in Reformer, recurrence in the Universal Transformer, and stabilization for reinforcement learning in GTrXL.

The post described the vanilla Transformer as an encoder-decoder architecture, citing Vaswani et al., 2017, noting later simplified versions such as encoder-only BERT and decoder-only GPT. It detailed the encoder's stack of six identity modules, each with a multi-head self-attention layer and a point-wise feed-forward network, and the decoder's two multi-head attention submodules. It described sinusoidal and learned positional encoding.

The post covered Adaptive Computation Time, citing Graves, 2016, as a mechanism for dynamically deciding computational steps in a recurrent neural network. It described Transformer-XL, citing Dai et al., 2019, as addressing context segmentation through hidden state reuse between segments and relative positional encoding. It described adaptive attention span, citing Sukhbaatar et al., 2019, and localized attention in the Image Transformer, citing Parmer et al., 2018.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

[Updated on 2023-01-27 : After almost three years, I did a big refactoring update of this post to incorporate a bunch of new Transformer models since 2020. The enhanced version of this post is here: The Transformer Family Version 2.0 . Please refer to that post on this topic.]