Keerthana Gopalakrishnan on why robotics is still in its GPT-2 era
Google DeepMind's Gemini Robotics lead says the viral running robots miss the point: robotics is still in its GPT-2 era, held back by contact, cross-embodiment, and errors that compound across a task, and she treats progress as an empirical question, not a prediction.
Running fast is the wrong benchmark
The viral robot Olympics showed humanoids outrunning the fastest humans, but foot speed was never the thing holding robots back. What is hard is contact with soft, shifting things, not speed.
I think we are very much focused on doing robots help people and do useful things in the physical world and I think I'm a very productive human a lot of my friends are also very productively employed but we don't run faster than Usain Bolt
Robotics is still in its GPT-2 era
Keerthana scores the whole field at a GPT-2 moment. Two things have to work before it reaches a GPT-3: few-shot learning that holds across many tasks, and one brain that transfers across different robot bodies.
my phone or your phone my computer Mac Linux it doesn't matter where you run it it kind of behaves you can expect this very similar behavior but here I think we are very subject to which robots that you act on
Two brains and a body
Gemini Robotics 2 is three models: ER2, a slower reasoning brain now exposed through the API; the action model that turns intent into motion from fingertips to feet; and a smaller on-device version that runs without the cloud.
Gemini robotics er you can think of it as like a system to brain that can do like very generic reasoning. It's based on the flash line of models but maybe more tuned towards robotics.
Build the generalist, then climb to mastery
A narrow robot can be excellent at one task, but every new task costs just as much as the first. A general baseline makes each new skill cheap to reach, the same path language models took.
the cost of doing the nth thing becomes exactly the same as the cost of doing the first or the second thing. But building a very general baseline then makes it much easy to quickly take it to tune it towards mastery.
Some mistakes you can take back, some you can't
Whether a robot can be trusted with a task depends partly on whether a mistake is recoverable, not just on average accuracy. Stacking Lego tolerates errors; cracking an egg on the floor does not.
Tasks that allow retries are much easier like if you are doing like Lego assembly it's fine if you make a mistake you can go back and fix it.
Errors multiply down the chain
Chaining a reasoning model and an action model across many steps means their success rates multiply, so even strong per-step accuracy collapses over a long task.
it is also a compounding of errors right like now if you have 10 tasks in sequence and now you have the ER model a certain X success rate and the VA model which has a Y success rate now the if you have a sequencing task you are basically compounding the error
Safety is a capability, and it starts with not falling over
Keerthana treats safety as part of what makes a robot useful, not a tax on it. In robots the first safety problem is not a rogue takeover, it is a heavy humanoid that falls because it is clumsy.
I think of safety as like a capability, right? Like I there is a lot of discussion around like safety and capabilities being at odds with each other, but people are not going to use an unsafe robot and unsafe agents.
No single data source will solve robotics
Teleoperation, sensor-based demos, and first-person human video each trade scale against precision, so Keerthana expects a mixture, not one winning source, and refuses to call it before the experiments do.
I don't think it's wise to take a very principled view about this because it all of this will change as new evidence emerges and better hardware emerges and you probably need all of all types of data.