The line between 'an AI that writes' and 'an AI that does' just got blurrier. Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that reads and writes text, images and video — and drives robots — inside one single neural network. If that sounds like a technical footnote, it isn't. Most AI you touch today is a relay race: your request gets handed to a text specialist, then an image specialist, then a video specialist, each with its own brain. Rho-1 throws away the baton and runs the whole track by itself.
What happened
One model, every format, no handoffs
In plain English, a research preview means the thing works and you can look at it, but it's not a finished product you'd build a company around yet. The model has 19 billion parameters — think of parameters as the dials and knobs a network tunes while it learns. More dials, more nuance, up to a point.
The interesting part isn't the size. It's the plumbing. Most multimodal systems — the ones that handle more than one type of data — quietly aren't one system at all. They're a traffic cop that glances at your request and routes it to the right specialist. Rho-1 skips the cop. Text, images, video and robot control actions all get turned into tokens — the same small chunks of data the model already knows how to handle — and everything sits in one shared context window. A context window is just the model's working memory: everything it can hold in mind at once.
The upshot? No tool calls, no external models, no middlemen. One network, one memory, all formats.
It generates video live — and doesn't flinch when you change your mind
Rho-1 generates continuous video in real time and responds to new instructions on the fly without restarting. If you've played with AI video tools, you know the drill: type a prompt, wait, get a clip, tweak the prompt, start over from scratch. Rho-1 keeps the camera rolling and adjusts mid-stream. It's the difference between talking to someone and mailing them a letter.
How it learned to move a robot without robot data
Robot training data is famously scarce — nobody has millions of hours of robots loading dishwashers. So Reka AI built an inverse dynamics model, which is a fancy way of saying it works backwards. The system watches ordinary internet videos and pulls out the control signals that must have produced whatever motion it sees. If the arm swung that way, the muscle commands were probably this. Those signals then drive the same weights that predict camera images. Same brain, different body.
This isn't Reka's first swing
Reka AI is no stranger to multimodal work. Back in April 2024 it shipped Reka Core, a multimodal language model that went head-to-head with GPT-4, Claude 3 and Gemini Ultra on benchmarks. Rho-1 also slots into the wider research push toward so-called world models — AI that builds an internal feel for how things move, change and behave.
One more number worth sitting with: Rho-1 trained on 320 H100 GPUs over about three months. H100s are Nvidia's top-shelf AI chips, and renting that many for a quarter isn't pocket change. This is a well-funded lab making a serious bet.
What it means for you
You don't need to care about architecture. You do need to care about what gets faster, cheaper and less fiddly. Here's the practical read, situation by situation.
At home
Right now, 'sort out my week' means three apps and a lot of retyping. You photograph a scribbled calendar, paste a voice memo into a notes app, then type the to-do list yourself. A model that takes all of it at once — image, audio, text — and hands back a plan in one pass is a different kind of assistant. Same logic applies to a fridge photo and a dinner question, or a broken appliance and a repair clip.
At work
Think about how much of your day is copying things between tools: a screen recording here, a summary there, a slide deck somewhere else. Single-context models collapse that relay into one request. Hand over the recording of a demo and ask for the summary, the slide outline and the list of issues — no exporting, no re-uploading, no re-explaining what you already showed. The tedious middle of knowledge work is exactly the part a shared context window deletes.
For business
Small teams get leverage that used to require headcount. A two-person shop could produce demo videos, answer support tickets stuffed with screenshots, and experiment with automating a physical arm — all from one model family. The strategic point isn't doing any single thing better. It's not needing an integration project between every pair of tools.
For studying
Picture a tutor that watches your handwritten solution in a video, corrects you mid-thought, and switches languages without losing the thread. That's the promise of one context window. Instead of 'describe your problem to the chatbot', you just show it.
For creators
Real-time video with live direction is a storyboard that animates while you talk. Say 'make it night', then 'hold on — closer', and it keeps going. For anyone who's ever burned an afternoon on re-renders, that's the headline, not the benchmark scores.
For extra income
The money sits in the gap between 'I can't make video' and 'I can make video'. Product clips, faceless channels, quick social ads, explainer animations for local businesses — every step that gets cheaper opens room for someone to sell the result. Robot arms, realistically, aren't your side hustle. Content is.
How to try it right now
Step 1: Start free. Before you chase a research preview, get comfortable with tools that mix formats — image in, text out; video in, plan out. mykreatool.com collects free AI tools you can run today without a credit card.
Step 2: Find the model. Go to Reka AI's site and look for the Rho-1 research preview. 'Research preview' is code for limited, gated access, so expect a sign-up or a waitlist rather than an instant API key.
Step 3: Pick one real task. Choose something you currently split across two or three tools — a meeting recording plus a written recap, say — and see whether one request can cover both.
Step 4: Test the no-restart trick. Give it an instruction, let it run, then change your mind halfway through. Adapting on the fly is the whole point, so that's the thing worth poking at.
Step 5: Don't rebuild your stack. Keep your current workflow running alongside it. Previews break, change and vanish.
Upsides and what changes
Fewer seams, fewer failures. Every time you hand data from one model to another, something gets lost in translation — a face, a mood, the last sentence of an instruction. One shared context window means fewer of those handoffs.
Video stops being a destination. Today, video is where you end up after exporting everything else. When a model reads and writes video natively, a clip becomes an input as ordinary as typing.
Robots get cheaper to teach. Pulling control signals out of everyday internet video is a genuine shortcut around the biggest bottleneck in robotics: nobody filmed the training data you need. If it holds up at scale, the floor drops for anyone building physical automation.
Interaction gets conversational. Real-time generation plus live course-correction makes these systems feel less like vending machines and more like collaborators.
Limitations
Keep your expectations calibrated. Rho-1 is a research preview, not a product with a price tag, a support team or an uptime guarantee — and Reka AI hasn't published independent benchmarks or third-party testing for it. Nineteen billion parameters is respectable, but it's not the enormous scale of the headline frontier models, so breadth across text, images, video and robot control may come with less depth in each than a dedicated specialist would deliver. The inverse dynamics trick, clever as it is, learns from videos that happen to exist, which isn't the same as data collected for the task, and robot control in research demos doesn't always survive contact with a messy real environment. Real-time video generation raises its own questions about quality, consistency and cost, and running anything like this yourself isn't happening — 320 H100s for three months is a data-center bill, not a laptop project. Add the usual privacy wrinkles: you're feeding video, screens and sometimes your home into a model, so read the terms before you hand over the family kitchen.
Conclusion and one action for today
Rho-1 is a bet that the future of AI isn't more specialists wired together, but one network that handles words, pictures, video and movement in the same breath. That's a big claim, and it's early. But the direction is hard to argue with: fewer handoffs, fewer seams, less copy-paste. The people who win here won't be the ones who understand tokens — they'll be the ones who stop splitting every task across five apps.
Your action for today: pick one job you currently do in two or three separate tools, and try running it as a single request in a multimodal AI tool. Do it with a free option so it costs you nothing but ten minutes — then join the Rho-1 research preview waitlist so you'll know the moment the real thing shows up.



Comments 0