High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Sergey Levine: evaluation

12 Sept 2025 Dwarkesh Podcast Fully autonomous robots are much closer than you think – Sergey Levine

“One theme here that is important to keep in mind is that the reason that those building blocks are so valuable is because the AI community has gotten a lot better at leveraging prior knowledge.”

— Sergey Levine

Source trail

Everything needed to verify it.

Speaker
Sergey Levine
Attribution
Verified speaker
Claim type
evaluation
Recorded
12 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…I find it super interesting that you're using the open-source Gemma model, which is Google's LLM that they released open source, and then adding this action expert on top. I find it super interesting that the progress in different areas of AI is based on not only the same techniques, but literally the same model. You can just use an open-source LLM and add this action expert on top. You naively might think that, “Oh, there's a separate area of research which is robotics, and there's a separate area of research called LLMs and natural language processing.” No, it's literally the same. The considerations are the same, the architectures are the same, even the weights are the same. I know you do more training on top of these open-source models, but I find that super interesting. One theme here that is important to keep in mind is that the reason that those building blocks are so valuable is because the AI community has gotten a lot better at leveraging prior knowledge. A lot of what we're getting from the pre-trained LLMs and VLMs is prior knowledge about the world. It's a little bit abstracted knowledge. You can identify objects, you can figure out roughly where things are in image, that sort of thing. But if I had to summarize in one sentence, the big benefit that recent innovations in AI give to robotics is the ability to leverage prior knowledge. The fact that the model is the same model, that's always been the case in deep learning. But it's that ability to pull in that prior knowledge, that abstract knowledge that can come from many different sources that's really powerful. I was talking to this researcher, Sander at GDM, and he works on video and audio models. He made the point that the reason, in his view, we aren't seeing that much transfer learning between different modalities. That is to say, training a language model on video and images doesn't seem to necessarily make it that much better at textual questions and tasks because images are represented at a different semantic level than text. His argument is that text has this high-level semantic representation within the model, whereas images and videos are just compressed pixels. When they're embedded, they don't represent some high-level semantic information. They're just compressed pixels. Therefore there's no transfer learning at the level at which they're going through the model. Obviously this is super relevant to the work you're doing. Your hope is that by training the model on the visual data that the robot sees, visual data generally maybe even from YouTube or whatever eventually, plus language information, plus action information from the robot itself, all of this together will make it generally robust. You had a really interesting blog post about why video models aren't as robust as language models. Sorry, this is not a super well-formed question. I just wanted to get a reaction.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence