AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Learning with not Enough Data Part 3: Data Generation

Collected Oct 1, 2026

Lilian Weng published Part 3 of her series on learning with not enough data, following Parts 1 and 2. The post covers two approaches for generating synthetic data for training: augmented data and new data.

Augmented data applies augmentation, distortion, and transformation to existing training samples while preserving key attributes. The post describes the goal of data augmentation as modifying input format, such as text wording or visual appearance, while semantic meaning stays unchanged.

For images, methods include basic operations like random cropping and resizing, color distortions, Gaussian blur, color jittering, horizontal flipping, and grayscale conversion. Task-specific strategies include AutoAugment (Cubuk et al. 2018), which frames learning optimal augmentation operations as an RL problem; RandAugment (Cubuk et al. 2019), which reduces the search space using a single magnitude parameter; Population Based Augmentation (Ho et al. 2019); and Unsupervised Data Augmentation (Xie et al. 2019). Image mixture methods include Mixup, Cutmix, and MoCHi.

Text augmentation covers lexical edits via Easy Data Augmentation (Wei & Zou 2019) with synonym replacement, random insertion, random swap, and random deletion, plus contextual augmentation and back-translation. Audio methods include audio mixup, time masking, frequency masking, and frequency shift. Architectural augmentation uses dropout masks and cutoff.

For new data, the post describes relying on pretrained models to generate data points when few or none exist, noting few-shot prompting as effective for language models. Language models can serve as noisy annotators, as explored by Wang et al. (2021) using GPT-3, or as data generators, via LAMBADA (Anaby-Tavor et al. 2019), Unsupervised Data Generation (Wang et al. 2021), and work by Han et al. (2021) on translation tasks.

The post also discusses quantifying generated data quality through affinity and diversity, as introduced by Gontijo-Lopes et al. (2020), and covers training with noisy data via regularization, robust learning objectives, label correction, and sample reweighting.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Here comes the Part 3 on learning with not enough data (Previous: Part 1 and Part 2 ). Let’s consider two approaches for generating synthetic data for training. Augmented data . Given a set of existing training samples, we can apply a variety of augmentation, distortion and transformation to derive new data points without losing the key attributes. We have covered a bunch of augmentation methods on text and images in a previous post on contrastive learning. For the sake of post completeness, I duplicate the section