Meta Reinforcement Learning
Meta Reinforcement Learning (Meta-RL) aims to train an agent that can solve unseen tasks quickly and efficiently by meta-learning over reinforcement learning tasks. The approach is described in a post that traces the topic's origins and outlines its formulation and algorithms.
The post states that the idea has roots in a 2001 paper by Hochreiter et al., which was proposed for supervised learning but has many resemblances to current Meta-RL methods. In modern deep learning, Wang et al. (2016) and Duan et al. (2017) simultaneously proposed a very similar idea, called RL^2 in the second paper. A Meta-RL model is trained over a distribution of MDPs and can learn to solve a new task quickly at test time.
Meta-RL is defined as doing meta-learning within reinforcement learning. Train and test tasks are usually different but drawn from the same family of problems; the post lists multi-armed bandits with different reward probabilities, mazes with different layouts, and the same robots with different physical parameters in simulation as examples. Each task is formulated as an MDP, and in RL^2 an extra finite-horizon parameter T is added to the MDP tuple. The test tasks are sampled from the same distribution or a slightly modified version.
According to the post, the main difference from ordinary RL is that the last reward and last action are incorporated into the policy observation in addition to the current state. Both Meta-RL and RL^2 implemented an LSTM policy, with hidden states serving as memory for tracking trajectory characteristics. The post describes three key components: a model with memory, a meta-learning algorithm, and a distribution of MDPs.
The post covers several meta-learning algorithms for Meta-RL. MAML (Finn et al. 2017) and Reptile (Nichol et al. 2018) update model parameters for generalization. Meta-gradient RL (Xu et al. 2018) treats hyperparameters such as the discount factor and bootstrapping parameter as meta-parameters tuned online. Evolved Policy Gradient (Houthooft et al. 2018) defines the policy gradient loss as a temporal convolution over past experience and uses Evolutionary Strategies. MAESN (Gupta et al. 2018) learns structured action noise. Episodic control, proposed by Lengyel and Dayan (2008), is described as keeping explicit records of past events and using them directly as a reference for new decisions; MFEC (Blundell et al. 2016) models memory as a table storing state-action pairs and Q-values.
The post also notes that Meta-RL matches ideas in the AI-GAs paper by Jeff Clune (2019), which involves meta-learning architectures, meta-learning algorithms, and automatically generated environments.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
In my earlier post on meta-learning , the problem is mainly defined in the context of few-shot classification. Here I would like to explore more into cases when we try to “meta-learn” Reinforcement Learning (RL) tasks by developing an agent that can solve unseen tasks fast and efficiently.