AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Domain Randomization for Sim2Real Transfer

Collected Oct 1, 2026

Lilian Weng published an article on domain randomization for sim2real transfer in robotics. The piece describes how deep RL algorithms are sample inefficient and collecting data on real robots is costly, so models are often trained in simulators that can provide effectively unlimited data. However, a reality gap between simulator and physical world, caused by inconsistent physical parameters such as friction, kp, damping, mass and density, plus incorrect physical modeling like collisions between soft surfaces, often causes failure on real robots.

The article lists system identification, domain adaptation and domain randomization as approaches to close the gap. System identification requires careful calibration, which is expensive, and physical parameters of the same machine may vary with temperature, humidity, positioning or wear. Domain adaptation uses transfer learning techniques, often built on adversarial loss or GANs, to match the real data distribution; it typically needs a decent amount of real data. Domain randomization instead creates varied simulated environments with randomized properties and trains a model to work across them, on the expectation that the real system is one sample in that distribution. DR may need little or no real data.

The article covers uniform domain randomization, which samples each bounded randomization parameter uniformly, and applies it to scene appearance and physical dynamics. It cites work at OpenAI Robotics in 2018, where visual and dynamics DR produced a policy working on a real dexterous robot hand for rotating an object to 50 successive random target orientations; early policies survived barely more than 5 seconds without dropping the object.

For why DR works, the article offers two non-exclusive explanations: DR as bilevel optimization (Vuong et al., 2019) and DR as meta-learning, citing the learning dexterity project (OpenAI, 2018), where an LSTM policy generalized across environmental dynamics while a feedforward policy without memory did not transfer to a physical robot. Guided domain randomization is presented with optimization for task performance, including AutoAugment (Cubuk et al., 2018), 'learning to simulate' (Ruiz, 2019), evolutionary methods using CMA-ES (Yu et al., 2019) and Meta-Sim (Kar et al., 2019). Matching real data distribution covers SimOpt (Chebotar et al., 2019) and RCAN (James et al., 2019). Guided by data in the simulator covers DeceptionNet (Zakharov et al., 2019) and Active domain randomization (Mehta et al., 2019).

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

In Robotics, one of the hardest problems is how to make your model transfer to the real world. Due to the sample inefficiency of deep RL algorithms and the cost of data collection on real robots, we often need to train models in a simulator which theoretically provides an infinite amount of data. However, the reality gap between the simulator and the physical world often leads to failure when working with physical robots. The gap is triggered by an inconsistency between physical parameters (i.e. friction, kp, dampi