A Brief Overview of Gender Bias in AI

The Gradient published an article reviewing a selection of research on gender bias in AI models and the challenges of measuring and mitigating it. The author defines bias as unequal, unfavorable, and unfair treatment of one group over another, and notes that AI research often treats gender as binary man/woman because it is easier to measure.
The article covers five studies. Bolukbasi et al. (2016) found gender bias in word embeddings trained on Google News articles, including the analogy man is to computer programmer as woman is to homemaker, and proposed a debiasing method based on gender-neutral words; the author notes the method would not work for modern Transformer-based systems.
Buolamwini and Gebru (2018) evaluated three commercial gender classifiers on four subgroups and found all performed better on male than female faces, better on lighter than darker faces, and worst on darker female faces, with error rates up to 34.7%, versus a maximum of 0.8% for lighter-skinned male faces. In response, Microsoft and IBM revised and expanded their training datasets. Rudinger et al. (2018) found coreference resolution models resolved male pronouns to occupations more than female or neutral pronouns; none predicted managers to be female, though the occupation is 38.5% female in the U.S. according to 2006 Census data.
Parrish et al. (2021) created the BBQ benchmark across nine social dimensions; tested models reinforced stereotypes 77% of the time in ambiguous contexts. Luccioni et al. (2023) found image-generation models under-represent marginalized identities; DALL-E 2 generated white men 97% of the time for prompts like CEO.
The author notes that as benchmarks are used more by model builders, models may be optimized only for the specific biases those benchmarks capture, leaving other biases unaddressed, and cites their own explorations of historical figures, occupational associations, and DALL-E 3 prompt transformations. They also raise cultural and geographic bias, noting most images in Open Images and ImageNet come from the US and Great Britain. Proposed mitigations cited include Datasheets for Datasets, Model Cards, and the AI Foundation Model Transparency Act of 2023. The author frames what it means to fix bias as a philosophical question, describing a cycle where biases are found, addressed in new model versions, and then rediscovered elsewhere.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
A brief overview and discussion on gender bias in AI