Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions within a single neural network.
According to the company, Rho-1 differs from most AI systems that route tasks to specialized models. It runs all modalities as tokens in one shared context window, without tool calls or external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.
Reka AI says the same weights that predict camera images also drive robot movements. To work around scarce robot training data, the company built an inverse dynamics model that pulls control signals from ordinary internet videos.
Rho-1 trained on 320 H100 GPUs over about three months.
Reka AI is not new to multimodal AI. In April 2024, the company shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. According to the report, the Rho-1 release fits a broader push in AI research toward so-called world models.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Reka AI's Rho-1 is a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single neural network. Trained on 320 H100 GPUs in about three months, it uses a fraction of the compute today's top models need. Instead of routing tasks to specialized systems, Rho-1 runs all modalities as tokens in one shared context window. The article Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model appeared first on The Decoder .