AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Thinking about High-Quality Human Data

Collected Oct 1, 2026

Lilian Weng has published a post titled "Thinking about High-Quality Human Data" examining how human annotation underpins modern deep learning model training. She notes that most task-specific labeled data comes from human annotation, including classification tasks and RLHF labeling for LLM alignment training, and acknowledges feedback and pointers from Ian Kivlichan.

The post describes two directions for approaching data quality: human raters relative to data quality, and data quality relative to model training. On the raters side, it covers task design, selecting and training a pool of raters, and collecting and aggregating data.

On crowdsourcing, the post cites a 1907 Nature paper on the "Vox populi" idea, in which an exhibition crowd's middlemost estimate of an ox's weight came very close to the true value. It also cites Callison-Burch (2009), an early study using Amazon Mechanical Turk for non-expert human evaluation on machine translation, including having non-experts create new gold reference translations; the post says correlation between experts' and crowdsourced translations was higher than between expert and machine translation outputs.

For rater agreement, the post lists majority voting, raw agreement, Cohen's Kappa, and probabilistic graph modeling, and describes MACE (Hovy et al. 2013), which estimates the likelihood of an annotator behaving as a spammer by providing random labels.

On disagreement, the post cites Aroyo and Welty (2015) on annotation "myths," and Rottger et al. (2021), who formulated descriptive and prescriptive paradigms for subjective NLP annotation. It cites Goyal et al. (2022) on annotator identity as a statistically significant factor in labeling identity-related content as toxic, and Wang et al. (2023), which compared Trust and Safety professional labels with crowd annotations and found agreement rates ranging from 0.96 on violence/gory to 0.25 on personal topics, with higher agreement on extreme and benign conversations.

The post also describes Zhang et al. (2023)'s taxonomy of rater-disagreement causes, disagreement deconvolution (Gordon et al. 2021), a multi-annotator model (Davani et al. 2021) tested on the Gab Hate Corpus, and jury learning (Gordon et al. 2022), which models individual annotators' labeling behavior and uses a Deep and Cross network, evaluated on a toxicity diversity dataset using mean absolute error.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

[Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally human data collection i