Introducing Robostral Navigate

Mistral AI announced Robostral Navigate, its first model built for embodied navigation. The 8B model takes RGB images and a plain-language instruction and moves a robot through an environment, such as leaving a lobby, walking through a corridor, entering a supply room, and stopping to face the second shelf.
The company said other models often employ depth sensors, LiDAR, or several cameras, while Robostral Navigate uses only one ordinary RGB camera and no depth sensors. It reports 76.6% success on R2R-CE (Room-to-Room in Continuous Environments) validation unseen, a benchmark for following instructions in environments held out of training. According to Mistral AI, this beats the best single-camera approach by 9.7 points and the best system using depth or multiple cameras by 4.5 points. The model also reports 79.4% success on validation seen.
Navigation is performed via pointing: given a task and a history of observations, the model predicts where the robot should move next by inferring image coordinates of the target location in the current camera view, along with the desired orientation upon arrival. When the target lies outside the current field of view, it falls back to displacements in the robot's local coordinate frame.
Mistral AI said the model is built entirely in-house without existing open-source VLMs, initialized from its vision-language model specialized for grounding tasks. It was trained entirely in simulation using a data generation pipeline that produced approximately 2.4 million trajectories across 350k scenes. A prefix-caching training algorithm with tree-based attention masking compresses an episode into a single sequence, reducing training tokens by 22×. Online reinforcement learning with CISPO improved the success rate by 3.2%.
Mistral AI said the model runs on wheeled, legged, and flying robots and generalizes across robot sizes. It aims to enable robots to navigate offices, homes, commercial buildings, and outdoor spaces.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Introducing Robostral Navigate: 8B model achieving 76.6% on R2R-CE with just a single RGB camera. No depth sensors, LiDAR, or multiple cameras needed.