Contrastive Representation Learning
Lilian Weng has published a technical overview of contrastive representation learning, describing its goal as learning an embedding space in which similar sample pairs stay close and dissimilar ones stay far apart. The overview states that contrastive learning applies to both supervised and unsupervised settings, and that on unsupervised data it is one of the most powerful approaches in self-supervised learning.
On training objectives, the overview notes that early contrastive loss functions involved only one positive and one negative sample, while the trend in recent objectives is to include multiple positive and negative pairs in one batch. It covers contrastive loss (Chopra et al. 2005), triplet loss from the FaceNet paper (Schroff et al. 2015), Lifted Structured Loss (Song et al. 2015), N-pair loss (Sohn 2016), NCE (Gutmann and Hyvarinen, 2010), InfoNCE from CPC (van den Oord et al. 2018), and Soft-Nearest Neighbors Loss (Salakhutdinov and Hinton 2007; Frosst et al. 2019).
The listed key ingredients are heavy data augmentation, large batch size and hard negative mining. The overview cites SimCLR experiments showing that composing random cropping with random color distortion is crucial for learning visual representations of images. It also describes sampling bias, referring to false negative samples, as capable of causing a significant performance drop, and cites Chuang et al. (2020) for a debiased loss and Robinson et al. (2021) for targeting hard negatives through modified sampling probabilities.
The overview also organizes methods for image and sentence embeddings, image augmentation strategies such as AutoAugment, RandAugment, PBA and UDA, feature clustering approaches including DeepCluster and SwAV, and supervised approaches including CLIP and Supervised Contrastive Learning.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The goal of contrastive representation learning is to learn such an embedding space in which similar sample pairs stay close to each other while dissimilar ones are far apart. Contrastive learning can be applied to both supervised and unsupervised settings. When working with unsupervised data, contrastive learning is one of the most powerful approaches in self-supervised learning .