Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands and Generates Video and Outputs Robot Actions in One
By Michal Sutter | October 5, 2026
Reka has released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch. A single neural network understands and generates text, images, and video, reasons over them, and outputs robot actions. Reka frames it as a direct replacement for agentic pipelines that pass work between modality-specific models.
What Rho-1 Changes
Most multimodal systems today are pipelines. A central model plans, then hands jobs to specialists for images, video, or detection. Each handoff adds latency, and each specialist sees only a narrow request.
Rho-1 removes those handoffs. Text, vision, and robotic actions become tokens inside one context window. According to Reka's research, one unedited session shows the full loop. The model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm, and explains the difference. All of that happens in five turns, with no tool call and no second model.
Key Takeaways for 2026
As agentic AI matures in 2026, the industry is moving away from chained specialist models toward unified architectures that reason across modalities in a single pass. Rho-1 exemplifies this trend: a 19B parameter model trained from scratch to handle text, images, video, and robot actions natively. This consolidation reduces latency, simplifies deployment, and opens new possibilities for real-time robotics and interactive media generation.
The research preview signals Reka's ambition to collapse the multimodal stack into one model that can perceive, reason, and act โ a significant step toward general-purpose embodied AI.
via MarkTechPost
