Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
12 Sept 2025 Dwarkesh Podcast Fully autonomous robots are much closer than you think – Sergey Levine
“I don't know if there was an episode like this in the training set, but just for fun I took one of the shorts and turned it inside out.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 12 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…If you look in a dictionary, they'll have the pronunciation of a word written in funny letters. That's basically International Phonetic Alphabet. It's an alphabet that is pretty much exclusively used for writing down pronunciations of individual words and dictionaries. You can ask an LLM to write you a recipe for making some meal in International Phonetic Alphabet, and it will do it. That's like, holy crap. That is definitely not something that it has ever seen because IPA is only ever used for writing down pronunciations of individual words. That's compositional generalization. It's putting together things you've seen in new ways. Arguably there's nothing profoundly new here because yes, you've seen different words written that way, but you've figured out that now you can compose the words in this other language the same way that you've composed words in English. That's actually where the emergent capabilities come from. Because of this, in principle, if we have a sufficient diversity of behaviors, the model should figure out that those behaviors can be composed in new ways as the situation calls for it. We've actually seen things even with our current models. In the grand scheme of things, looking back five years from now, we'll probably think that these are tiny in scale. But we've already seen what I would call emerging capabilities. When we were playing around with some of our laundry folding policies, we actually discovered this by accident. The robot accidentally picked up two T-shirts out of the bin instead of one. It starts folding the first one, the other one gets in the way, picks up the other one, throws it back in the bin. We didn't know it would do that. Holy crap. Then we tried to play around with it, and yep, it does that every time. It's doing its work. Drop something else on the table, it just picks it up and puts it back. Okay, that's cool. It starts putting things in a shopping bag. The shopping bag tips over, it picks it back up, and stands it upright. We didn't tell anybody to collect data for that. I'm sure somebody accidentally at some point, or maybe intentionally picked up the shopping bag. You just have this kind of compositionality that emerges when you do learning at scale. That's really where all these remarkable capabilities come from. Now you put that together with language. You put that together with all sorts of chain-of-thought reasoning, and there's a lot of potential for the model to compose things in new ways. Right. I had an example like this when I got a tour of the robots at your office. It was folding shorts. I don't know if there was an episode like this in the training set, but just for fun I took one of the shorts and turned it inside out. Then it was able to understand that it first needed to get… First of all, the grippers are just like this, two opposable finger and thumb-like things. It's actually shocking how much you can do with just that. But it understood that it first needed to fold it inside out before folding it correctly. What's especially surprising about that is it seems like this model only has one second of context. Language models can often see the entire codebase. They're observing hundreds of thousands of tokens and thinking about them before outputting. They're observing their own chain of thought for thousands of tokens before making a plan about how to code something up. Your model is seeing one image, what happened in the last second, and it vaguely knows it's supposed to fold this short. It's seeing the image of what happened in the last second. I guess it works. It's crazy that it will just see the last thing that happened and then keep executing on the plan. Fold it inside out, then fold it correctly. But it's shocking that a second of context is enough to execute on a minute-long task. Yeah. I'm curious why you made that choice in the first place and why it's possible to actually do tasks… If a human only had a second of memory and had to do physical work, I feel like that would just be impossible. It's not that there's something good about having less memory, to be clear. Adding memory, adding longer context, all that stuff, adding higher resolution images, those things will make the model better. But the reason why it's not the most important thing for the kind of skills that you saw when you visited us, at some level, comes back to Moravec's paradox. Moravec's paradox basically, if you want to know one thing about robotics, that's the thing. Moravec's paradox says that in AI the easy things are hard and the hard things are easy. Meaning the things that we take for granted—like picking up objects, seeing, perceiving the world, all that stuff—those are all the hard problems in AI. The things that we find challenging, like playing chess and doing calculus, actually are often the easier problems. I think this memory stuff is actually Moravec’s paradox in disguise. We think that the cognitively demanding tasks that we do that we find hard, that cause us to think, "Oh man, I'm sweating. I'm working hard." Those are the ones that require us to keep lots of stuff in memory, lots of stuff in our minds. If you're solving some big math problem, if you're having a complicated technical conversation on a podcast, those are things where you have to keep all those puzzle pieces in your head. If you're doing a well-rehearsed task—if you are an Olympic swimmer and you're swimming with perfect form—and you're right there in the zone, people even say it's "in the moment." It's in the moment. It's like you've practiced it so much you've baked it into your neural network in your brain. You don't have to think carefully about keeping all that context. It really is just Moravec's paradox manifesting itself. That doesn't mean that we don't need the memory. It just means that if we want to match the level of dexterity and physical proficiency that people have, there's other things we should get right first and then gradually go up that stack into the more cognitively demanding areas, into reasoning, into context, into planning, all that kind of stuff. That stuff will be important too.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.