Self-Supervised Representation Learning
Lilian Weng's post on self-supervised representation learning surveys methods that generate labels from data itself via pretext tasks. The motivating idea, as stated in the post, is that manually labeling data such as ImageNet is expensive and hard to scale, while unlabeled data is plentiful; self-supervised learning frames a supervised task to predict a subset of information using the rest, so inputs and labels are both provided by the data. The learned intermediate representation is expected to carry semantic or structural meaning useful for downstream tasks, rather than the pretext task's own accuracy mattering.
For images, the post describes a workflow of training on pretext tasks with unlabeled images, then feeding an intermediate feature layer to a multinomial logistic regression classifier on ImageNet classification to quantify representation quality. It notes some researchers train supervised learning on labeled data and self-supervised pretext tasks on unlabeled data simultaneously with shared weights, citing Zhai et al., 2019 and Sun et al., 2019.
Image-based approaches covered include distortion, such as Exemplar-CNN (Dosovitskiy et al., 2015) using 32x32 patches, and rotation prediction (Gidaris et al., 2018) as a 4-class problem. Patches methods include relative position prediction (Doersch et al., 2015), which samples a first patch and a second from one of 8 neighboring locations in a 3x3 grid, and jigsaw puzzles (Noroozi & Favaro, 2016) that place 9 shuffled patches back. Feature counting (Noroozi et al., 2017) uses scaling and tiling relations with an MSE loss plus a difference loss.
Colorization (Zhang et al., 2016) maps grayscale images to a distribution over quantized color values in CIE Lab* space, with a rebalanced loss. Generative modeling includes denoising autoencoders (Vincent et al., 2008), context encoders (Pathak et al., 2016), split-brain autoencoders (Zhang et al., 2017) predicting color channels, and bidirectional GANs (Donahue et al., 2017). Contrastive learning covers Contrastive Predictive Coding (van den Oord et al., 2018), whose InfoNCE loss uses cross-entropy to classify future representations among unrelated negative samples.
Update notes list additions: Contrastive Predictive Coding on 2020-01-09; a Momentum Contrast section on MoCo, SimCLR and CURL on 2020-04-13; a Bisimulation section on DeepMDP and DBC on 2020-07-08; MoCo V2 and BYOL on 2020-09-12; and on 2021-05-31 removal of the Momentum Contrast section with a pointer to a full post on Contrastive Representation Learning.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
[Updated on 2020-01-09: add a new section on Contrastive Predictive Coding ]. [Updated on 2020-04-13: add a “Momentum Contrast” section on MoCo, SimCLR and CURL.] [Updated on 2020-07-08: add a “Bisimulation” section on DeepMDP and DBC.] [Updated on 2020-09-12: add MoCo V2 and BYOL in the “Momentum Contrast” section.] [Updated on 2021-05-31: remove section on “Momentum Contrast” and add a pointer to a full post on “Contrastive Representation Learning” ]