Reka’s Rho-1: One AI Model Trying to Do the Work of Many

Published:

Most AI assistants that seem to “do everything” are really a team of specialists behind the scenes. One model plans, then hands the job to a separate image model, a video model or a tool. Every hand-off adds delay, and each specialist only sees part of the picture.

AI company Reka says it has built something different. On 5 October 2026 it released a research preview of Rho-1, which it describes as a 19B “omni-reasoning” model, trained from scratch, that understands and generates text, images and video, reasons over them and takes actions, all inside one neural network [1]. “19B” refers to about 19 billion parameters, the internal settings a model learns during training; that makes it mid-sized by today’s standards.

What does Rho-1 do?

Reka’s demo shows a single conversation in which the model draws a picture of a lighthouse, puts a box around the lighthouse, turns the picture into a short video, edits the video into a snowstorm, then explains in words what changed [1]. Reka says every step came from the same model, with no tool calls and no second model.

The company pitches three uses [1]:

  • A multimodal assistant that can move between words, pictures and video in one session.
  • A steerable “world model”: a simulated scene that keeps running and changes when you give new instructions mid-stream, which Reka suggests could be used to test self-driving or robot software.
  • A robotics brain that predicts what a camera will see next and outputs the movements to make it happen. Reka says it can pair Rho-1 with its own Inverse Dynamics Model, a system that works out the likely control movements from ordinary video, so web video can be used as training data.

How does it work, in plain terms?

Rho-1 treats everything as tokens, the small units a model reads and writes. Text and commands are discrete tokens, like words. Images, video frames and robot actions are continuous tokens, which keep their natural, smooth form. Both kinds share the same attention, the mechanism that lets a model relate one piece of input to another [1].

Reka says it trains the model with two methods at once: next-token prediction for text and flow matching, a technique for generating images, video and actions [1].

How fast is it?

Reka says the base model generates video at 0.79 times real time (median), and that a faster “distilled” version cuts the clean-up steps used in generation from 99 to 8 with minimal quality loss [1]. In its own internal testing, the company says the distilled model was among the fastest in every type of output it measured.

It also stresses the modest budget: Reka says the model was trained from scratch on 320 H100 GPUs (specialised AI chips) for three months, a small fraction of the computing power behind frontier models [1].

What are the limitations?

Reka is unusually upfront about these. It lists [1]:

  • Long-horizon drift: in longer clips the layout of a scene can drift even if it still looks realistic.
  • Grounding across time: locating objects works on still images but not yet reliably across video.
  • Editing stability: targeted edits are brittle across different prompts.
  • Resolution: native video is capped at 672×384.

Reka calls Rho-1 a proof of concept and an architectural direction, not a finished product [1].

What we don’t know

  • Independent results. The speed and capability comparisons come from Reka’s internal tests, not a public leaderboard or outside evaluation.
  • Real-world readiness. It is a research preview with acknowledged failure modes, and Reka has not said when, or whether, it will become a product.
  • Whether scaling fixes the problems. Reka believes more compute and data will narrow its limitations; that is its expectation, not a demonstrated result.

The Bottom Line

Rho-1 is an interesting bet that one mid-sized model holding everything in a single “memory” can replace a chain of specialist systems. The demos are striking and the training budget is small. But until outsiders can test it, the speed claims remain Reka’s own, and the company itself says the hard problems are not solved yet.

Sources

  1. Reka, 5 October 2026: announcement of the Rho-1 research preview. https://reka.ai/news/rho-1-collapsing-the-multimodal-stack
TSN
TSNhttps://tsnmedia.org/
Welcome to TSN. I'm a data analyst who spent two decades mastering traditional analytics—then went all-in on AI. Here you'll find practical implementation guides, career transition advice, and the news that actually matters for deploying AI in enterprise. No hype. Just what works.

Related articles

Recent articles