Evidence receipt / belief
Published · transcript-backedSergey Levine: belief
12 Sept 2025 Dwarkesh Podcast Fully autonomous robots are much closer than you think – Sergey Levine
“" Representing your context in the right form, that captures what you really need to achieve your goal—and otherwise discards all the unnecessary stuff—I think that's a really important thing.”
Source trail
Everything needed to verify it.
- Speaker
- Sergey Levine
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 12 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Exactly. You have to have many order-of-magnitude improvements across all of these three things which seem to oppose each other. Increasing one reduces the amount of compute you can dedicate towards the other one in inference. How are we going to solve this? That's a very big question. Let's try to unpack this a little bit. There's a lot going on in there. One thing is a really interesting technical problem. It's something where we'll see perhaps a lot of really interesting innovation over the next few years. It’s the question of representation for context. You gave some of the examples, like if you have a home robot that's doing something then it needs to keep track. As a person, there are certainly some things where you keep track of them very symbolically, almost in language. I have my checklist. I'm going shopping. At least for me, I can literally visualize in my mind my checklist. Pick up the yogurt, pick up the milk, pick up whatever. I'm not picturing the milk shelf with the milk sitting there. I'm just thinking, “milk.” But then there's other things that are much more spatial, almost visual. When I was trying to get to your studio, I was thinking, "Okay, here's what the street looks like. Here's what that street looks like. Here's what I expect the doorway to look like. " Representing your context in the right form, that captures what you really need to achieve your goal—and otherwise discards all the unnecessary stuff—I think that's a really important thing. We're seeing the beginnings of that with multimodal models. But I think that multimodality has much more to it than just image plus text. That's a place where there's a lot of room for really exciting innovation. Do you mean in terms of how we represent?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.