High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Sergey Levine: evaluation

12 Sept 2025 Dwarkesh Podcast Fully autonomous robots are much closer than you think – Sergey Levine

“Those things can also be addressed with scaling. But we have to identify the right axes for that, which means figuring out what data to collect, what settings to collect it in, what methods consume that data, and how those methods work.”

— Sergey Levine

Source trail

Everything needed to verify it.

Speaker
Sergey Levine
Attribution
Verified speaker
Claim type
evaluation
Recorded
12 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…What is preventing you now from scaling that data even more? If data is a big bottleneck, why can't you just increase the size of your office 100x, have 100x more operators operating these robots and collecting more data. Why not ramp it up immediately 100x more? That's a really good question. The challenge here is understanding which axes of scale contribute to which axes of capability. If we want to expand capability horizontally—meaning the robot knows how to do 10 things now and I'd like it to do 100 things later—that can be addressed by just directly horizontally scaling what we already have. But we want to get robots to a level of capability where they can do practically useful things in the real world. That requires expanding along other axes too. It requires, for example, getting to very high robustness. It requires getting them to perform tasks very efficiently, quickly. It requires them to recognize edge cases and respond intelligently. Those things can also be addressed with scaling. But we have to identify the right axes for that, which means figuring out what data to collect, what settings to collect it in, what methods consume that data, and how those methods work. Answering those questions more thoroughly will give us greater clarity on the axes, on those dependent variables, on the things that we need to scale. We don't fully know right now what that will look like. I think we'll figure it out pretty soon. It's something we're working on actively. We want to really get that right so that when we do scale it up, it'll directly translate into capabilities that are very relevant to practical use. Just to give an order of magnitude, how does the amount of data you have collected compare to internet-scale pre-training data? I know it's hard to do a token-by-token count, because how does video information compare to internet information, et cetera. But using your reasonable estimates, what fraction?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence