Reducing Toxicity in Language Models
Lilian Weng has published a post on reducing toxicity in language models, noting that large pretrained models trained on online data unavoidably acquire toxic behavior and biases from the Internet. The post states that safely deploying them for real-world applications demands strong safety control over the model generation process.
Among the challenges the post lists are the variety of unsafe content types, including toxicity, abusiveness, hate speech, biases, stereotypes, cyberbullying and identity attacks, and the lack of a widely agreed-upon categorization and definition of unsafe behavior in pretrained language models, since individual perceptions can vary by social background. The post cites definitions from Perspective API, Kurita et al. 2019 and Pavlopoulos et al. 2020.
On categorization, the post describes the three-level hierarchical taxonomy proposed by Zampieri et al. 2019, which considers both the type and the target of offense, and is the basis for the Offensive Language Identification Dataset (OLID).
For data collection, the post cites Vidgen and Derczynski 2020 on annotation methods including expert coding, crowdsourcing, professional moderators and synthetic data, noting crowdsourcing is the most common. It describes Khatri et al.'s 2018 two-stage bootstrap classifier and SOLID, which contains 9+ million tweets and extends OLID via democratic co-training.
Detection topics include adversarial attacks, with Dinan et al.'s 2019 build-break-fix strategy and Kurita et al.'s character-level perturbations, Perspective API scoring, and Schick et al.'s 2021 self-diagnosis. Detoxification covers blacklisting, vocabulary shifting, and prompt-based self-debiasing. The post also mentions a system-level safety solution.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior and biases from the Internet. Pretrained language models are very powerful and have shown great success in many NLP tasks. However, to safely deploy them for practical real-world applications demands a strong safety control over the model generation process.