AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Learning with not Enough Data Part 1: Semi-Supervised Learning

Collected Oct 1, 2026

Lilian Weng published Part 1 of a planned series titled "Learning with not Enough Data," focused on semi-supervised learning. The post describes semi-supervised learning as training a model on labeled and unlabeled data together.

Four approaches are listed for limited labeled data: pre-training plus fine-tuning, semi-supervised learning, active learning, and pre-training plus dataset auto-generation. The post notes that most existing semi-supervised learning literature addresses vision tasks, while pre-training plus fine-tuning is a more common paradigm for language tasks.

The described methods combine a supervised loss with an unsupervised loss, often weighted by a ramp function that increases the unsupervised term's importance over the training step.

Four hypotheses are cited: smoothness assumptions, cluster assumptions, low-density separation assumptions, and manifold assumptions.

Consistency regularization methods covered include the Pi-Model, credited to Laine and Aila (2017) after Sajjadi et al. (2016), which minimizes differences between two passes through a network with stochastic transformations; Temporal Ensembling, which maintains an exponential moving average of predictions per sample; and Mean Teacher (Tarvaninen and Valpola, 2017), which instead averages model weights.

Noisy-sample methods include Adversarial Training (Goodfellow et al. 2014), Virtual Adversarial Training (Miyato et al. 2018), Interpolation Consistency Training (Verma et al. 2019), and Unsupervised Data Augmentation (Xie et al. 2020).

Pseudo Labeling (Lee 2013) assigns fake labels based on maximum softmax probabilities and is described as equivalent to Entropy Regularization (Grandvalet and Bengio 2004). Label Propagation (Iscen et al. 2019) diffuses labels over a similarity graph.

Self-Training (Scudder 1965; Nigram and Ghani, CIKM 2000) iterates between training on labeled data and converting the most confident predictions into labels. Noisy Student used an EfficientNet teacher to pseudo-label 300M unlabeled images and a larger student trained with noise. SentAugment (Du et al. 2020) retrieves in-domain unlabeled sentences via sentence embeddings for self-training in language.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

When facing a limited amount of labeled data for supervised learning tasks, four approaches are commonly discussed.