Skip to content
Zenteck
Latest
Analysis

Reka Rho-1: Omnimodal AI Processes Text, Video, and Robots in One Network

Por Alex da Cruz2 min read0 comments
reka-rho-1-ia-unifies-text-video-and-robos-on-a-unique-network
A

Alex da Cruz

Alex da Cruz is a full-stack developer based in São Paulo, Brazil. He works with React, TypeScript and automation, and uses AI daily to solve real problems in code and operations — not as a demo. He has run an e-commerce operation end to end, and now builds and maintains the automation pipeline behind this blog. He writes about what he actually tests.

Ver perfil →

According to The Decoder, Reka AI launched a research preview of Rho-1, a 19-billion-parameter model designed to process multiple types of data simultaneously. Instead of routing complex requests to specialized separate models, the system runs text, images, video, and robotic actions as shared tokens within a single context window.

How does Rho-1 handle robot control without specialized data?

Training physical robots usually requires scarce, specialized motion capture datasets. To bypass this bottleneck, Reka AI utilized an inverse dynamics model that extracts control signals directly from standard internet videos. The identical model weights that predict camera pixels are also used to drive physical robot movements.

What does this mean for automation workflows?

Running all modalities through a single neural network removes the latency typically introduced by tool calls and external model handoffs. For developers building automation pipelines, this unified architecture means video generation and real-time instructions can happen concurrently without restarting the system or orchestrating multiple APIs.

According to the report, the model was trained on 320 H100 GPUs over approximately three months, demonstrating a fraction of the compute resources typically required by larger frontier labs for multimodal systems.

Sources

  1. Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model — the-decoder.com

Frequently asked questions

What is Reka Rho-1?
Rho-1 is a 19-billion-parameter omnimodal model by Reka AI that processes text, images, video, and robot control inside a single neural network.
How many GPUs were used to train Rho-1?
According to The Decoder, the model was trained on 320 H100 GPUs over a period of about three months.