Learning with not Enough Data Part 2: Active Learning
This is part two of a series on learning with limited labeled data for supervised tasks. This installment addresses active learning, where a limited budget for human labeling requires selecting which samples to label.
Given an unlabeled dataset and a fixed labeling cost, active learning aims to select a subset of examples to be labeled that can maximize improvement in model performance. The post focuses on deep neural models and batch-mode training, assuming a K-class classification problem in which a model with parameters outputs a probability distribution over labels.
The scoring function used to identify valuable examples is called an acquisition function. Basic sampling strategies include uncertainty sampling, which selects examples with the most uncertain predictions using scores such as least confident, margin, and entropy, and query-by-committee, which uses a committee of models with measures including voter entropy, consensus entropy, and KL divergence.
Diversity sampling seeks a collection of samples that represent the entire data distribution. Expected model change refers to a sample's impact on model training. Hybrid strategies combine preferences, such as selecting uncertain but representative samples.
The post covers deep acquisition functions for measuring uncertainty, including aleatoric uncertainty from data noise and epistemic uncertainty within model parameters. Ensemble approaches include MC dropout, which approximates a probabilistic deep Gaussian process, and naive ensembles, which performed better than cheaper alternatives in cited comparisons. Other methods discussed include Bayes-by-backprop, loss prediction, adversarial setups such as VAAL, MAL, and CAL, and representativeness measures including core-sets.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
This is part 2 of what to do when facing a limited amount of labeled data for supervised learning tasks. This time we will get some amount of human labeling work involved, but within a budget limit, and therefore we need to be smart when selecting which samples to label.