Curriculum for Reinforcement Learning
An article on curriculum learning for reinforcement learning presents several categories of curricula, including Task-Specific Curriculum, Teacher-Guided Curriculum, Curriculum through Self-Play, Automatic Goal Generation, Skill-Based Curriculum, and Curriculum through Distillation. The article was updated on 2020-02-03 to mention PCG in the Task-Specific Curriculum section, and updated on 2020-02-04 to add a new curriculum through distillation section.
The article states that the idea of training neural networks with a curriculum was proposed by Jeffrey Elman in 1993. His early work on learning simple language grammar demonstrated the importance of starting with a restricted set of simple data and gradually increasing the complexity of training samples; otherwise the model was not able to learn at all, according to the article. It notes that a curriculum may expedite convergence and may or may not improve final model performance, and that a bad curriculum may hamper learning.
The article describes Bengio et al. (2009) as providing an overview of curriculum learning and presenting two ideas through toy experiments with a manually designed task-specific curriculum: cleaner examples may yield better generalization faster, and introducing gradually more difficult examples speeds up online training. It also cites Weinshall et al. (2018) on quantifying task difficulty using minimal loss with respect to a pretrained model, and Zaremba and Sutskever (2014) on training an LSTM to predict outputs of short Python programs without executing the code, where a combined curriculum reportedly always outperformed a naive curriculum and generally, but not always, outperformed a mix strategy.
The article covers teacher-guided curriculum learning, including Automatic Curriculum Learning proposed by Graves et al. in 2017 and formalized as Teacher-Student Curriculum Learning (Matiisen et al., 2017), where a student works on actual tasks while a teacher selects sub-tasks. It also discusses asymmetric self-play (Sukhbaatar et al., 2017), automatic goal generation with Goal GAN (Florensa et al., 2018), and later work by Racaniere and Lampinen et al. (2019) on more sophisticated goal generators. The article attributes the observation that uniformly sampling from all tasks is a surprisingly strong benchmark to the two discrete task-space works it cites.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
[Updated on 2020-02-03: mentioning PCG in the “Task-Specific Curriculum” section. [Updated on 2020-02-04: Add a new “curriculum through distillation” section.