TextPurr Logo

TextPurr

Loading...
Loading...

Google DeepMind's Kanishka Rao on the Road to Physical AGI

The Humanoid Hub
My conversation with Kanishka Rao, a director on Google DeepMind's robotics team, who went from a particle physics PhD at CERN to driving robotics at DeepMind. We get into the Gemini Robotics 2 release, trending training recipes in the industry, physical AGI, world models, and more. 1:00 Intro 1:47 Does AI's pace feel crazy? 4:04 From particle physics to robotics 6:03 Speech recognition's deep learning moment 8:43 Whole-body humanoid control 10:47 The trash-bag tying demo 13:07 The robotics data pyramid 14:13 Transferring skills across robots 16:30 How does Motion Transfer work 16:56 Why a separate reasoning model 19:02 Can distillation work in robotics? 20:02 The ASIMOV safety benchmark 22:47 VLAs vs world models 25:10 The missing tactile sensing piece 27:48 How far to physical AGI? 30:07 Fundamental breakthrough needed? 31:53 World models for evaluation 33:45 Will the model layer commoditize? 35:11 Humanoid as a form factor Recorded on: 27th July, 2026 Hosted by: Devang Adhyaru Follow The Humanoid Hub on X: https://x.com/thehumanoidhub Subscribe to weekly newsletter: https://thehumanoidhub.com/get-newsletter
Hosts: Kanishka Rao, Devang
📅July 30, 2026
⏱️00:38:27
🌐English

Disclaimer: The transcript on this page is for the YouTube video titled "Google DeepMind's Kanishka Rao on the Road to Physical AGI" from "The Humanoid Hub". All rights to the original content belong to their respective owners. This transcript is provided for educational, research, and informational purposes only. This website is not affiliated with or endorsed by the original content creators or platforms.

Watch the original video here: https://www.youtube.com/watch?v=CrudsEeyS6k

00:00:00Kanishka Rao

So yeah, I did my PhD in particle physics, and I got really lucky. I got to work at CERN during the time where they discovered the Higgs boson. I got to see that firsthand.

💬 0 comments
00:00:09Kanishka Rao

A trash bag is not like a rigid object, so it constantly, you know, gets... it's easy to slip and miss. So we thought to ourselves, if you can solve this task, then we kind of know we have an unlock here.

💬 0 comments
00:00:17Kanishka Rao

That's a great question, and to be honest, this question keeps me up at night.

💬 0 comments
00:00:22Kanishka Rao

We actually explored this a lot in our recent release of Gemini Robotics 2. We really focused on unlocking the humanoid form.

💬 0 comments
00:00:32Kanishka Rao

And I think it was mostly, I think maybe for a couple of reasons, by putting like a wrist camera next to your contact points, and then that maybe gets you out of the hole. But I think ultimately we will have to face this tactile thing in robotics, and dexterity will get really good up until a point, and then I think it will break down because you don't have tactile information.

💬 0 comments
00:01:00Devang

Gemini Robotics has got legs, and literally so. Google DeepMind has been doing frontier robotics for years, and its latest release takes things to another level. Robots can walk, balance, and use their hands with real dexterity, all driven by a single AI model.

💬 0 comments
00:01:16Devang

So here to talk about that is Kanishka Rao. He is a director on the robotics team at Google DeepMind, leading the efforts to bring AI out of the lab and into real robots. Kanishka, thanks for joining.

💬 0 comments
00:01:31Kanishka Rao

Thanks so much for having me.

💬 0 comments
00:01:37Devang

So to start with, before we get into the news, I have to ask this. There's a lot going on in the AI space lately, in the last three years. It feels like we are living in a surreal time, and you've got a front-row seat, not just developing but observing all this. Does it feel a little bit crazy at times to you?

💬 0 comments
00:02:03Kanishka Rao

Yes. I think I never kind of get used to that feeling, that this is kind of normal. So I completely agree with you. Yeah, it's been a crazy journey in the past few years—not only in robotics, but just, you know, as you were saying, just generally how AI is developing every year.

💬 0 comments
00:02:18Kanishka Rao

This year alone, we've seen all these agentic workflows with these harnesses, so it's really cool to see, one, how quickly AI as a whole, like the whole community, is improving in terms of capabilities, and then also how this translates to the robotic space. I'm obsessed with the physical AGI part of it, so just seeing how all these breakthroughs—those same kind of lessons—all seem to transfer to robotics.

💬 0 comments
00:02:44Kanishka Rao

In the end, our goal is to essentially just ride the same wave of these AI breakthroughs but in the robotic space, and leverage all those breakthroughs to make robots more useful. So yeah, it feels uncomfortably exciting how quickly things are progressing.

💬 0 comments
00:03:03Devang

And people get used to it so soon. I'm sure you do get used to it. Robots were dancing two years ago, and now nobody cares. It must be happening a lot, and if you take a step back, it does look amazing in hindsight.

💬 0 comments
00:03:21Kanishka Rao

Yeah, I try really hard, or maybe it comes naturally. I generally, whenever I walk into the lab, have this small feeling of, "This is so cool." Then, of course, the mundane day-to-day stuff takes over, but I try to really take a second every day to appreciate how far we've come.

💬 0 comments
00:03:34Kanishka Rao

I've been working on robotics in this team close to eight years now. I just think back to year eight, year seven, year six, and it's just incredible, the pace of improvements. Then you start thinking in the future, like eight years from now. It's easy to get kind of used to it, but in robotics, you see the robots moving every day, so it's in your face more, maybe. But yeah, it's really incredible to see the improvements.

💬 0 comments
00:04:04Devang

Yeah. So can you tell us a bit about your journey? You've done a PhD in physics and astronomy, which is quite different from what you're doing right now. So how did the inspiration come for going into software and robotics, and how did you end up here?

💬 0 comments
00:04:23Kanishka Rao

Yeah, my journey was kind of all over the place, not really well-planned. Actually, up until a point, I think I really wanted to do physics. Growing up, that was my one passion. I really wanted to learn how the universe works and some of those questions. So yeah, I did my PhD in particle physics, and I got really lucky. I got to work at CERN during the time where they discovered the Higgs boson. I got to see that firsthand, so I got incredibly lucky there.

💬 0 comments
00:04:47Kanishka Rao

I got to see how these big organizations work. If you have a huge, ambitious goal, how can people from all over the world come to solve that problem? The scale of that organization, CERN, was really impressive, as was the rigor at which they do science.

💬 0 comments
00:05:02Kanishka Rao

Really, I was pursuing physics, and one of my graduate school friends was interviewing at Google. He told me Google was hiring physics PhDs. I was very curious why Google would want a physics PhD, so I decided to interview at Google. It was a software engineering interview, so I had to study what a binary tree was—I had no idea.

💬 0 comments
00:05:26Kanishka Rao

Somehow, I passed those interviews and got an offer. I had my postdoc in physics in one hand and this weird software engineering role from Google in the other. I decided at that point, "Hey, I'll just try Google for a year just to get a sense of what that world is like, and then I'll come back and do my postdoc after that." I think that was 13 years ago, and I'm still here.

💬 0 comments
00:05:49Devang

Wow.

💬 0 comments
00:05:49Kanishka Rao

I think it's been really satisfying doing research at Google and DeepMind because it feels like you're still at the cutting edge and you're still exploring. Some of those same itches that I got from physics kind of got met here.

💬 0 comments
00:05:59Kanishka Rao

I worked on speech recognition for five years. Again, I got very lucky in my career. When I first joined speech, it was very much like a research field, and the speech recognition error rates were too high for it to be usable. But the week I joined, deep neural networks—specifically RNNs—became very popular, and all of a sudden, the speech errors started plummeting.

💬 0 comments
00:06:31Kanishka Rao

Very quickly, a year after I joined, we had these things called LSTMs that were very, very good. All of a sudden, speech went from this research field to, "Hey, this technology is now usable and useful." I got to see it become a product. Then, one day, we had 100 million users using it every day. It was really cool to see how research breakthroughs can impact and benefit people, becoming a product used all over the world.

💬 0 comments
00:06:57Kanishka Rao

Seeing that really fueled me to ask, "What is the next big unsolved thing?" That's what brought me to robotics about eight years ago. I'm still kind of waiting, looking for that moment where robotics is useful enough and solved enough that it's going to go outside into the real world and start impacting and benefiting people at scale. So that's currently how I got to robotics.

💬 0 comments
00:07:23Devang

Do you wonder if the two fields will collide at some point, where the research in particle physics may provide some research avenues in robotics?

💬 0 comments
00:07:36Kanishka Rao

Yeah, I hope so. I think even Demis talks about this a lot. One of the aspirational or ultimate goals for AGI is to enhance or solve science and to basically increase human knowledge of our universe. So yeah, I'm hopeful that some of these AGI-based scientists can come up with new theories. There are so many fundamental questions in physics that are still out there, things like dark matter, quantum gravity, or unsolved problems. I'm hoping that some of these AIs, when they're advanced enough, can do this.

💬 0 comments
00:08:11Kanishka Rao

And then also the physical AGI part: all of these experiments in physics require massive detectors to be built. It is physical work, and there's inspecting the detectors when there's a lot of radiation in there where humans can't go. So there, I'm hoping that robots can help build these detectors or maintain them. Both the digital and the physical AIs will help advance physics knowledge at some point, and that will be really incredible to see.

💬 0 comments
00:08:41Devang

Yeah. Let's get into the new stuff from Google DeepMind, which is Gemini Robotics 2, the shiny new release in the robotics policy and control space. The new thing that caught my eye is that it can now control the whole body of a humanoid robot—it can walk, it can crouch—and it uses the same model, the Gemini Robotics 2 model. Why does it matter that the same model can control the walking part of the humanoid? Because usually what you see is that part is trained by sim-to-real and then fused into a VLA using that controller.

💬 0 comments
00:09:24Kanishka Rao

I think even in our release, there's still a controller that is in charge of balancing. What we have here is the visual-language-action model providing targets for that walking. Basically, I think this is important because we're combining whole-body control with the manipulation part of it. If you just have a dancing robot, for example, then you just have your controller that dances. But if you want to go pick up something from the floor, you have to be very particular about where you step, how far you are from that object when you start crouching, and how you extend your arm.

💬 0 comments
00:09:57Kanishka Rao

So we're really focused on the manipulation and interaction part of this whole-body control, where you cannot really separate the legs from the hands. How you step and how you move your hips all matter for you to then pick up that object from the back of the shelf when you're crouching. It's really critical that all these joints are controlled by a single model, just as a human would think about how they move.

💬 0 comments
00:10:20Devang

I see. So the VLA is kind of agnostic about which kind of robot it is?

💬 0 comments
00:10:24Kanishka Rao

Well, to some extent. It still knows, for example, how many joints are in there because it gives actions that are suitable for that robot. So I think it knows some of the embodiment prompts, as you would say, but it's not built for a specific embodiment. It's built for any embodiment, but it still has some embodiment-specific knowledge to control those specific joints.

💬 0 comments
00:10:45Devang

I see. And the really impressive demo was the robot with five-finger hands tying drawstrings of a trash bag. That was done with the Apollo robot from Apptronik, right? So what kind of unlocked this dexterous manipulation? Was it specifically planned into the architecture, or did this emerge as a capability?

💬 0 comments
00:11:16Kanishka Rao

Yeah, I think we've been pursuing dexterity for a while, and we've really been limited by the hardware, right? Initially, in previous releases, we talked about a single parallel gripper that essentially has $1\text{ DoF}$. Then I think we've also used, in previous times, hands that have had five fingers, but you cannot control them with that many degrees of freedom.

💬 0 comments
00:11:40Devang

Right.

💬 0 comments
00:11:40Kanishka Rao

With this release, we are controlling these Shadow hands, and each hand has 22 degrees of freedom. So it's starting to look closer to all the flexibility and range of motion that a human has. It's not just grasping, but you can control each one of these fingers.

💬 0 comments
00:11:58Kanishka Rao

I think the breakthrough here was having a single model able to control all these degrees of freedom together with your vision to do that task. That trash bag task was the most difficult task that we could pick so that we could solve the research for it. It's our show-off task because it's very, very difficult. It has to do very precise moves and deal with very flimsy materials. The plastic of a trash bag is not like a rigid object, so it constantly gets easy to slip and miss.

💬 0 comments
00:12:28Kanishka Rao

So we thought to ourselves, if we can solve this task, then we kind of know we have an unlock here, as opposed to a much easier task. We had to push on the boundaries of what the hardware is capable of, and that led to the modeling breakthroughs of how we get the model to control all these degrees of freedom to come together to do this task.

💬 0 comments
00:12:44Kanishka Rao

The nice thing is that once you unlock this recipe for hands, then you can do a lot more tasks with it. Even though we developed it with a few tasks in mind, the methods themselves are general. We can use them now to solve many, many more tasks that go beyond simple hardware to do much more human-like tasks.

💬 0 comments
00:13:06Devang

Right. And for that kind of dexterous manipulation with five-finger hands, do you use teleoperation data, or is there a way to scale egocentric—just human egocentric—data to train this kind of behavior?

💬 0 comments
00:13:19Kanishka Rao

Yeah, so I think the recipe to do this is not fully solved yet, and we iterate on this with every single release. But what I can say is it is a mix of all the data sources. I think us and a bunch of other people talk about this data pyramid in robotics. We have teleoperation data, we have simulation data, we have human egocentric data. All of these have a role to play. What mix they have and which one is useful for which kind of intelligence is still a recipe we're iterating on; it's not fully solved. But I think we're close.

💬 0 comments
00:13:54Devang

Mm-hmm.

💬 0 comments
00:13:54Kanishka Rao

We'll know which ways, because it's not only about which data is useful, but also which one is scalable. These all scale in different ways. Really, to solve any manipulation task in the world, you want lots and lots of data. So we really think about the scaling laws of each one of these and then trade off between them.

💬 0 comments
00:14:13Devang

Right. And even though there are five-finger robotic hands, all the hands are different in certain ways. If you are training teleop on one robot, does it really transfer to other robot types, and does the AI learn from all of them together?

💬 0 comments
00:14:30Kanishka Rao

Yeah, we're seeing signs of this, and they're very encouraging. I think that's one of the benefits of how we're approaching robotics: we don't develop a brain for a single embodiment. We try to have a nice representative set of robots in our model. We've introduced more complicated humanoid robots, but we still have gripper-based platforms that are more fixed tabletop things.

💬 0 comments
00:14:56Kanishka Rao

In our previous release, Gemini Robotics 1.5, we talked about motion transfer techniques where data collected on one robot showed signs of life on another robot. We still continue to see some of this in our portfolio because, since we have such a diverse portfolio of robots, the model learns to think about tasks in an embodiment-agnostic way. Because it's seen so many different forms, it learns that picking up an apple with a hand looks like this, with a gripper looks like this, and with a head looks like that.

💬 0 comments
00:15:33Kanishka Rao

Just by the scaling laws, you want to have a lot of diversity in the physical form, and that encourages the model to abstract away and just learn the essence of what a task is. So yeah, we're seeing some of those signs of life that we can then generalize to new embodiments.

💬 0 comments
00:15:43Kanishka Rao

In our release, we took our models trained on some of our more data-heavy platforms and then fine-tuned them on robots we have very little data on. Very quickly, the model is able to adapt. Within hours of new data, it can connect the dots and think, "I've seen all this other data with all these different embodiments. The only thing different here is the embodiment. I can quickly figure out somehow that I need to change these many things to get me to be performant with this embodiment."

💬 0 comments
00:16:15Kanishka Rao

We're seeing some very encouraging signs that this is a good strategy to train a foundational model for robotics, because you want to have a nice representative set of embodiments so that your model can deal with an unseen embodiment.

💬 0 comments
00:16:30Devang

You mentioned motion transfer. Is there a separate module which is retargeting all the motions, or is it a learned algorithm?

💬 0 comments
00:16:37Kanishka Rao

I can't talk about the specifics, but it's essentially a learned thing. Again, our philosophy is definitely that data, learning, and scale are what will solve everything. We definitely don't do anything that's engineered or specific in that way. We definitely bet on scale and deep learning to solve the day.

💬 0 comments
00:16:56Devang

Another key piece of the ecosystem is the embodied reasoning model. That one understands the instructions, plans, and reasons around the scene. Why is that a separate model and not part of the larger model that does everything in a single pass?

💬 0 comments
00:17:15Kanishka Rao

That's a great question. To be honest, personally, I've gone back and forth between these two. Talking about this bet on scaling, another kind of bet is, "No, it should be a single model, end-to-end trained, and it will just do everything in the end." But we're seeing that there are some very powerful intermediate steps.

💬 0 comments
00:17:33Kanishka Rao

If you take today, for example, the coding agents that we're all used to and use so much, even there, we're seeing this notion of a harness being so powerful. A good harness can make a mediocre model look good and a good model really excel. I think similar reasons make this ER model very attractive.

💬 0 comments
00:17:56Kanishka Rao

In our latest release, we are using it as a little bit of a harness where it's able to communicate with the user, orchestrate the VLAs, and then understand the progress of the VLA to decide when to switch tasks. This combination of a separate ER model that's really good at one thing and a VLA that's really good at another thing is working really well.

💬 0 comments
00:18:22Kanishka Rao

As I said earlier in the interview, we're really trying to see how the digital AI space is playing out and learn lessons from them. This explosion of autonomous agents with these harnesses inspired us to create this kind of agent now. The ER model is more of an agentic model that can basically use robots as physical agents to do some of these much more collaborative and long-horizon tasks. At least for the next few years, it seems very reasonable that you'll have some of these two- or three-model systems that really excel at doing some of these elaborate tasks.

💬 0 comments
00:19:00Devang

In the digital AI space, distillation seems to be working pretty powerfully. Do you think those techniques can be applied in robotics with multiple models?

💬 0 comments
00:19:11Kanishka Rao

I think it's hard to predict that because physical distillation will be quite different. If I'm understanding your question correctly, it would require unrolling the model on a physical robot and getting the trajectories that way, or in simulation.

💬 0 comments
00:19:28Devang

Yeah, or in simulation.

💬 0 comments
00:19:28Kanishka Rao

In simulation, if you can access enough of the state space there, you may be able to unroll the models enough to get that. A lot of the boundaries between good models and bad models are failure cases, so your simulation would have to be pretty good at getting those counterfactuals, like, "Hey, if I'm picking up a cup, if I hit it and knock it over, it should spill." Those boundary conditions are going to be very critical to capture. But yeah, potentially some of these distillation techniques might also work for physical models.

💬 0 comments
00:20:02Devang

You introduced a safety benchmark. I haven't seen many safety benchmarks in the robotics industry. What inspired getting this safety benchmark out? In terms of choosing between safety and performance, where do you draw the line? Because if the system is being too safe or too conservative, you miss out on doing some capable tasks.

💬 0 comments
00:20:30Kanishka Rao

Yeah, that's a good question. Robotics is very, very competitive right now, but we've kind of agreed in our team, at least, that safety should not be one of those things. I think safety should be more collaborative across robotics labs. So we're trying to basically put out more and share more there.

💬 0 comments
00:20:45Kanishka Rao

This benchmark is one step towards that, where we want to explicitly show how we're measuring our own safety. The ASIMOV benchmark is a great one where we've talked about safety in terms of reasoning. In this update, since we've made it more agentic, we have more agentic safety use cases. We really hope that this will showcase not only all the challenges associated with robotic safety, but also provide a clear benchmark where people can compare their various safety systems. We really are more open about our safety benchmarks and how we track progression there.

💬 0 comments
00:21:23Kanishka Rao

To your question about capability and safety, that's a great one too. We really have to make sure that the intelligence that powers safety has to be on par with the intelligence that powers capabilities. If one of them is off, you're going to get a mismatch. If safety is too powerful, then it will just stop doing any actions. If it's too weak, then it'll be too dangerous. We build the safety system over the same things that the reasoning and the action systems are built on—the same kind of Gemini models—so at least there's an intelligence match between them.

💬 0 comments
00:22:06Kanishka Rao

In these early stages, it's okay to be overprotective and overly careful about some of these as we're still understanding their capabilities. We have a pretty extensive agentic safety layer now that controls the entire stack.

💬 0 comments
00:22:22Kanishka Rao

Really, robotic safety is going to be the most important thing to solve. People are very focused on capabilities right now, but if you want to deploy these to homes, factories, and logistics spaces, safety will be the biggest thing to solve. You'll have to think about it all the way from the agentic orchestration layer all the way down to controls and the physical collision layer of the safety problem.

💬 0 comments
00:22:50Devang

Switching to a little bit of longer-term trends in the industry. Jim Fan at Nvidia recently declared there is an end of the VLA paradigm, and he and others are betting on world action models instead. There are different camps, with companies like Physical Intelligence and Sanctuary training models from scratch using UMI data, and there are other approaches as well. Do you find any of those other approaches that are not VLAs interesting, and do you expect the industry to converge on a certain recipe? Because underneath all of it is just the transformer algorithm.

💬 0 comments
00:23:37Kanishka Rao

Yeah, great question. Even though we've made so much progress and we talk about these step changes, I still feel that, going back to that speech moment where it felt solved, robotics is not solved today. There are still, I believe, some fundamental breakthroughs needed to get to that place where it starts feeling like it's ready to be deployed, useful, and benefit people.

💬 0 comments
00:23:52Kanishka Rao

These questions of what kind of model is the backbone intelligence of actions—is it a VLM, is it a world model?—or even the data question of, "Hey, is it human data collection versus teleoperation?" We're actively exploring all of these, and we've been working on them for years, so we definitely have been exploring these questions.

💬 0 comments
00:24:27Kanishka Rao

But I think it's just unsolved yet, so it's hard to claim what's going to work. I do think that these trends go through cycles where a breakthrough is made and everybody exploits it. Then you get to the end of the exploitation and there's more exploration. We'll go through these cycles. Some of these ideas will work out, and then people will start converging on them, finding the limitations of that, and then we'll have another exploration phase.

💬 0 comments
00:24:55Kanishka Rao

I feel like over the next few years, some of these ideas will make it or break it. Maybe after two or three years of these cycles, we'll be close to getting to the final answer of how robotics is roughly solved with a general recipe.

💬 0 comments
00:25:10Devang

Most of the datasets we see are vision-based with some robot joint action data, but we don't see much sensory data for hands. Is that going to change the way people think about these architectures?

💬 0 comments
00:25:24Kanishka Rao

That's a great question, and to be honest, this question keeps me up at night. I really care about how dexterous robots can get and how they use their hands. You've pointed out a big, almost flaw in current robot models: essentially all the state-of-the-art models today, like Gemini Robotics 2, are vision-based. You can have proprioceptive input, but there's no tactile information.

💬 0 comments
00:25:54Kanishka Rao

But if you think about all the manipulation that we humans do day-to-day, it's mostly tactile-based, right? You're not really looking very carefully. When you're tying a trash bag, you're not looking at every single way your finger is moving; you're kind of just going by feel or memory. There's this mismatch between how we think humans solve these dexterous tasks day-to-day and how robots are solving them.

💬 0 comments
00:26:16Kanishka Rao

I think there are a couple of reasons for that. One is the bet on using frontier models to solve robotics. All these frontier models are trained on the internet of data, which has tons and tons of visuals and images, so you have a very strong prior of what the world looks like. You have almost no dataset in terms of tactile information to touch that even compares.

💬 0 comments
00:26:41Kanishka Rao

One way this could play out is that vision is just going to be king because frontier models are built off vision, and you can leverage that by putting a wrist camera next to your contact points, and then that maybe gets you out of the hole. But I think ultimately we will have to face this tactile thing in robotics. Dexterity will get really good up until a point, and then I think it will break down because you don't have tactile information.

💬 0 comments
00:27:08Kanishka Rao

Another reason why we really haven't been able to do this research well is because skin is a really difficult hardware problem. Replicating all the richness that comes out of your skin sensors, especially on your palms, is so rich and diverse that it's going to be a while before we've figured out how to do and replicate that.

💬 0 comments
00:27:31Kanishka Rao

I think there will come a time when skin will get solved, hopefully, and we'll be able to include tactile data into our frontier models. But until then, I think vision will take us a long way, and it'll be interesting to see how far it actually goes.

💬 0 comments
00:27:48Devang

How far do you think we are from physical AGI, if we can even define what physical AGI is? And what do you think are the key limiting factors to getting there?

💬 0 comments
00:28:00Kanishka Rao

I think physical AGI is very easy to define, unlike digital AGI. Physical AGI, at least to me, is: I have a robot in front of me—let's say it's a humanoid—and I ask it to do anything that I could do, and it should be able to do it. And that's basically it.

💬 0 comments
00:28:17Kanishka Rao

I don't think it requires superhuman capabilities like a specialist human would have. Whatever an average human adult can do, the robot should be able to do it in the same time. I think time is also critical; you can't take an hour to do something that takes me five seconds. So that's the definition of physical AGI for me, and I think we're kind of far from there.

💬 0 comments
00:28:39Kanishka Rao

Some of the big breakthroughs needed are around the usual bets in robotics: generalization, mastery, dexterity, and understanding. Humans are very good at being thrown into any new scene and being able to do the task. All the breakthroughs we're making are still not at the human level. Once we start measuring against humans and getting to that bar, then we'll start hitting physical AGI. The definition for me is very human-centric, and that's the bar. We're still not there, as we talked about with how dexterous humans can get, how high our success rates are at doing things, and how quickly we can do them. A few breakthroughs are still needed for us to get to that point.

💬 0 comments
00:29:31Devang

Does it matter how fast the AI is able to learn? Because if someone is new at a job, they take anywhere between five to ten days to get a physical task going. Does it really matter if the AI does it in months or days?

💬 0 comments
00:29:46Kanishka Rao

No, I think that's a great point. A physical AGI definition must include that you can learn a new task as quickly as humans can. You need some kind of in-context learning or something quickly on the job that you can absorb and replicate. It's a great point.

💬 0 comments
00:30:04Devang

You mentioned... I was wondering if we have all the pieces of the puzzle to solve physical AGI, or is there a fundamental breakthrough needed in either computing or algorithms for all of them to converge and enable physical AGI?

💬 0 comments
00:30:25Kanishka Rao

That's a great question, and it's hard to say. Especially on continual learning, as you just pointed out, I think there are still algorithmic breakthroughs needed. Maybe these algorithms already exist today, and probably all the ideas that need to be solved for robotics are out there today. It's not so much about generating new ideas as it is about understanding which ideas make sense here, executing them really well, and showing that they will scale.

💬 0 comments
00:30:54Kanishka Rao

I don't think there's some fundamental computing breakthrough needed. Given the level of intelligence that we've seen from digital AIs today, it seems like robotics does not require much more sophisticated computational intelligence to solve these physical tasks. I think it's just a matter of data bottlenecks, understanding some of these algorithmic differences, and unlocking those.

💬 0 comments
00:31:30Kanishka Rao

As much as digital AI progresses, robotics has hitched its wagon to that. Those step changes also take robotics with them. Maybe it's just a matter of waiting until those models get so good that robotics gets solved for free on top of them.

💬 0 comments
00:31:55Devang

Google DeepMind also has some amazing generative world models. Has the robotics team been using some of those for either evaluation or synthetic data generation?

💬 0 comments
00:32:07Kanishka Rao

Yeah, we even published some work showing that we can use them for evaluation. The really attractive part about some of those models is that they seem to understand motion quite well. If I have a prompt that says, "Pour coffee into a mug," they do quite well. At I/O, Demis talked about how there seems to be some physics understanding in these models, and that's very attractive from a robotics perspective. If you can build a perfect physics simulator of the world, then you basically solve robotics because you can learn that understanding from it.

💬 0 comments
00:32:44Kanishka Rao

Going back to your point of language models versus motion models, I think both have their trade-offs. With language-based VLAs, we were able to build instruction-following robots that you can talk to and understand their thought processes. Some of these motion models are very attractive because they might be able to tell us how to understand motions better and generalize better in those motions. We're looking into some of these properties, and it's very exciting to see where this kind of research will lead us.

💬 0 comments
00:33:20Devang

And are world models going to become good at fine things like dexterity and understanding those movements?

💬 0 comments
00:33:27Kanishka Rao

It's hard to predict, but it's a matter of what's in the data and what's in your benchmarks. Maybe robotics can help improve those aspects in the world model, and it becomes a feedback loop.

💬 0 comments
00:33:43Devang

Yeah, potentially.

💬 0 comments
00:33:46Devang

I wanted to ask about something that a lot of developers in robotics are wondering. On the frontier digital AI side, there is a popular opinion among tech leaders that the model layer will become commoditized. We are seeing that in programming AI a little bit right now. Do you think the infrastructure part and the application part in robotics will be more sustainable economics-wise compared to the model, or is it a whole systems problem?

💬 0 comments
00:34:27Kanishka Rao

Yeah, very good question. Again, it's hard to predict that, but I feel like with physical AI, there is a pretty large overhead for deployment. You do need to make sure that your robot has the correct harness and the correct infrastructure to hook up the model.

💬 0 comments
00:34:46Kanishka Rao

Even though we talk a lot about how well cross-embodiment works, there's still work to be done today to fine-tune for a particular embodiment and figure out its quirks. So I'm not sure which layer will be a commodity, but definitely the AI part seems like the bottleneck right now in terms of capabilities.

💬 0 comments
00:35:11Devang

And how does Google DeepMind think about humanoids? Is it just another form factor that we need to solve, or should we treat that as a preferred general-purpose form and something that other robots can also benefit from?

💬 0 comments
00:35:27Kanishka Rao

We actually explored this a lot in our recent release of Gemini Robotics 2. We really focused on unlocking the humanoid form. I think it was mostly for a couple of reasons. One is that it really advanced the research. There are some research assumptions you might make on a gripper-based, simple tabletop robot, like we've done in previous releases, that may not hold for a much more complicated robot.

💬 0 comments
00:35:50Kanishka Rao

Focusing on the humanoid form with whole-body intelligence and high degrees of freedom hands really pushed us to revisit some of those assumptions and make sure that our intelligence still works for the humanoid. It remains a great embodiment to really push the state-of-the-art of a frontier model.

💬 0 comments
00:36:11Kanishka Rao

Again, I believe physical AGI would mean that whatever I can do, the robot should be able to do. That leads us down a path of something that has my kind of form to be able to replicate that. So I think that is critical.

💬 0 comments
00:36:25Kanishka Rao

The second reason is we definitely want to make sure we don't just build models for humanoids; it should be a general-purpose model that works for all robots. That forces us to ensure we don't take shortcuts for the humanoid or hardcode anything that's humanoid-specific. It really forced us to make sure all our modeling methods were agnostic to the form.

💬 0 comments
00:36:50Kanishka Rao

I think it's good for pushing research, and it's good because our world is built for humans and we want to make sure we can do all the tasks that humans can do. But on the other side, we also want to make sure that our models are completely humanoid-agnostic because of another trap: humanoid robots will keep changing. Whatever form we're training on today, five years from now, it will be different with different joints and configurations. We really don't want to anchor to any single embodiment; we want the AI to be embodiment-proof. Keeping that portfolio of embodiments keeps us honest to that endeavor.

💬 0 comments
00:37:31Devang

Right. To close it off, I wanted to know: is there a specific task at your home that you would like a robot to do?

💬 0 comments
00:37:41Kanishka Rao

So many! That's one of my biggest motivations for working on robotics: I really dread doing a lot of home chores. If a robot can do that in my home, then I'll be very happy. A lot of the tasks that we all dread—doing the dishes, doing my laundry, taking out the trash, cleaning the house—anything I could think of. All of those would be great if a robot can do those for us in the future.

💬 0 comments
00:38:11Devang

Yeah, that would be awesome. That was it from me. Did you have anything that we might have missed?

💬 0 comments
00:38:19Kanishka Rao

No, it was a pleasure talking to you, Devang. Thank you for having me on.

💬 0 comments
00:38:23Devang

Yeah, thank you so much. It was a pleasure. Thanks, bye.

💬 0 comments
Video Player