Evidence receipt / evaluation
Published · transcript-backedSergey Levine: evaluation
12 Sept 2025 Dwarkesh Podcast Fully autonomous robots are much closer than you think – Sergey Levine
“Meaning the things that we take for granted—like picking up objects, seeing, perceiving the world, all that stuff—those are all the hard problems in AI.”
Source trail
Everything needed to verify it.
- Speaker
- Sergey Levine
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 12 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Right. I had an example like this when I got a tour of the robots at your office. It was folding shorts. I don't know if there was an episode like this in the training set, but just for fun I took one of the shorts and turned it inside out. Then it was able to understand that it first needed to get… First of all, the grippers are just like this, two opposable finger and thumb-like things. It's actually shocking how much you can do with just that. But it understood that it first needed to fold it inside out before folding it correctly. What's especially surprising about that is it seems like this model only has one second of context. Language models can often see the entire codebase. They're observing hundreds of thousands of tokens and thinking about them before outputting. They're observing their own chain of thought for thousands of tokens before making a plan about how to code something up. Your model is seeing one image, what happened in the last second, and it vaguely knows it's supposed to fold this short. It's seeing the image of what happened in the last second. I guess it works. It's crazy that it will just see the last thing that happened and then keep executing on the plan. Fold it inside out, then fold it correctly. But it's shocking that a second of context is enough to execute on a minute-long task. Yeah. I'm curious why you made that choice in the first place and why it's possible to actually do tasks… If a human only had a second of memory and had to do physical work, I feel like that would just be impossible. It's not that there's something good about having less memory, to be clear. Adding memory, adding longer context, all that stuff, adding higher resolution images, those things will make the model better. But the reason why it's not the most important thing for the kind of skills that you saw when you visited us, at some level, comes back to Moravec's paradox. Moravec's paradox basically, if you want to know one thing about robotics, that's the thing. Moravec's paradox says that in AI the easy things are hard and the hard things are easy. Meaning the things that we take for granted—like picking up objects, seeing, perceiving the world, all that stuff—those are all the hard problems in AI. The things that we find challenging, like playing chess and doing calculus, actually are often the easier problems. I think this memory stuff is actually Moravec’s paradox in disguise. We think that the cognitively demanding tasks that we do that we find hard, that cause us to think, "Oh man, I'm sweating. I'm working hard." Those are the ones that require us to keep lots of stuff in memory, lots of stuff in our minds. If you're solving some big math problem, if you're having a complicated technical conversation on a podcast, those are things where you have to keep all those puzzle pieces in your head. If you're doing a well-rehearsed task—if you are an Olympic swimmer and you're swimming with perfect form—and you're right there in the zone, people even say it's "in the moment." It's in the moment. It's like you've practiced it so much you've baked it into your neural network in your brain. You don't have to think carefully about keeping all that context. It really is just Moravec's paradox manifesting itself. That doesn't mean that we don't need the memory. It just means that if we want to match the level of dexterity and physical proficiency that people have, there's other things we should get right first and then gradually go up that stack into the more cognitively demanding areas, into reasoning, into context, into planning, all that kind of stuff. That stuff will be important too. You have this trilemma. You have three different things which all take more compute during inference that you want to increase at the same time. You have the inference speed. Humans are processing 24 frames a second or whatever it is. We can react to things extremely fast. Then you have the context length. For the kind of robot which is just cleaning up your house, I think it has to be aware of things that happened minutes ago or hours ago and how that influences its plan about the next task it's doing. Then you have the model size. At least with LLMs, we've seen that there's gains from increasing the amount of parameters. I think currently you have 100 millisecond inference speeds. You have a second-long context and then the model is a couple billion parameters? Each of these, at least two of them, are many orders of magnitude smaller than what seems to be the human equivalent. A human brain has trillions of parameters and this has like 2 billion parameters. Humans are processing at least as fast as this model, actually a decent bit faster, and we have hours of context. It depends on how you define human context, but hours of context, minutes of context.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.