AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Exploration Strategies in Deep Reinforcement Learning

Collected Oct 1, 2026

Exploration versus exploitation is described as a critical topic in reinforcement learning, where agents are expected to find the best solution quickly while avoiding premature commitment that could lead to local minima or failure. Modern RL algorithms optimize returns and achieve exploitation efficiently, but exploration remains an open topic. The post was updated on 2020-06-17 to add "exploration via disagreement" in the "Forward Dynamics" section.

Classic exploration strategies covered include epsilon-greedy, upper confidence bounds, Boltzmann exploration, and Thompson sampling. For deep RL, the post mentions entropy loss terms and noise-based exploration, citing Fortunato, et al. 2017 and Plappert, et al. 2017.

Two key exploration problems are outlined. The hard-exploration problem refers to environments with very sparse or deceptive rewards, with Montezuma's Revenge cited as a concrete example. The noisy-TV problem, described as a thought experiment in Burda, et al. 2018, involves an agent rewarded by novel experience that becomes attracted to an uncontrollable random-noise TV, failing to make meaningful progress.

Intrinsic rewards as exploration bonuses are presented through a combined reward r_t = r^e_t + beta r^i_t, where beta is a hyperparameter. Count-based exploration is discussed via pseudo-counts derived from density models, with Bellemare, et al. 2016 using a CTS density model and Georg Ostrovski, et al. 2017 improving it with PixelCNN. Zhao and Tresp 2018 used a variational GMM. Hashing approaches are attributed to Tang et al. 2017, using locality-sensitive hashing and SimHash.

Prediction-based exploration is traced to Schmidhuber 1991. Forward dynamics models map (s_t, a_t) to s_{t+1}. IAC (Oudeyer, et al. 2007) uses prediction error and learning progress. Stadie et al. 2015 normalized prediction error. ICM (Pathak, et al. 2017) learns state encoding via an inverse dynamics model. Burda, Edwards and Pathak, et al. 2018 compared encoding functions for curiosity-driven learning.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

[Updated on 2020-06-17: Add “exploration via disagreement” in the “Forward Dynamics” section . Exploitation versus exploration is a critical topic in Reinforcement Learning. We’d like the RL agent to find the best solution as fast as possible. However, in the meantime, committing to solutions too quickly without enough exploration sounds pretty bad, as it could lead to local minima or total failure. Modern RL algorithms that optimize for the best returns can achieve good exploitation q