Reka Rho-1: Omnimodal AI Processes Text, Video, and Robots in One Network

Alex da Cruz
Alex da Cruz is a full-stack developer based in São Paulo, Brazil. He works with React, TypeScript and automation, and uses AI daily to solve real problems in code and operations — not as a demo. He has run an e-commerce operation end to end, and now builds and maintains the automation pipeline behind this blog. He writes about what he actually tests.
According to The Decoder, Reka AI launched a research preview of Rho-1, a 19-billion-parameter model designed to process multiple types of data simultaneously. Instead of routing complex requests to specialized separate models, the system runs text, images, video, and robotic actions as shared tokens within a single context window.
How does Rho-1 handle robot control without specialized data?
Training physical robots usually requires scarce, specialized motion capture datasets. To bypass this bottleneck, Reka AI utilized an inverse dynamics model that extracts control signals directly from standard internet videos. The identical model weights that predict camera pixels are also used to drive physical robot movements.
What does this mean for automation workflows?
Running all modalities through a single neural network removes the latency typically introduced by tool calls and external model handoffs. For developers building automation pipelines, this unified architecture means video generation and real-time instructions can happen concurrently without restarting the system or orchestrating multiple APIs.
According to the report, the model was trained on 320 H100 GPUs over approximately three months, demonstrating a fraction of the compute resources typically required by larger frontier labs for multimodal systems.
Sources
Frequently asked questions
- What is Reka Rho-1?
- Rho-1 is a 19-billion-parameter omnimodal model by Reka AI that processes text, images, video, and robot control inside a single neural network.
- How many GPUs were used to train Rho-1?
- According to The Decoder, the model was trained on 320 H100 GPUs over a period of about three months.
Comments
0 comments
Be the first to comment.
Continue Lendo

US Tech Giants Agree on Voluntary AI Oversight Framework
US tech giants signed a voluntary AI oversight framework with the government, outlining four pillars for future development.

FTC AI probe: what changes for companies using models
The FTC probe into major AI labs introduces new regulatory and compliance risks for companies building workflows on third-party models.

Meta Muse AI Privacy: Why Permission Layers Matter
Meta disputes claims that its Muse AI read private messages without permission, pointing to strict macOS security layers.