Evidence receipt / uncertainty
Published · transcript-backedLeopold Aschenbrenner: uncertainty
4 Jun 2024 Dwarkesh Podcast Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history
“I don't know how rocket science works, but I didn't get the impression that there's some clear scaling behavior with the amount of jet fuel.”
Source trail
Everything needed to verify it.
- Speaker
- Leopold Aschenbrenner
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 4 Jun 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Let’s go back to a more fundamental thing we're talking about here. In the series you say we should denominate the probability of getting to AGI in terms of OOMs of effective compute. Effective here accounts for the fact that there's a compute multiplier if you have a better algorithm. I'm not sure that it makes sense to be confident that this is a sensible way to project progress. It might be, but I have a lot of uncertainty about it. It seems similar to somebody trying to project when we're going to get to the moon. They're looking at the Apollo program in the 1950s or something. They're like, "we have some amount of effective jet fuel and if we get more efficient engines, then we have more effective jet fuel. So we're going to determine the probability of getting to the moon based on the amount of effective jet fuel we have." I don't deny that jet fuel is important to launch rockets, but that seems like an odd way to denominate when you're going to get to the moon. I don't know how rocket science works, but I didn't get the impression that there's some clear scaling behavior with the amount of jet fuel. First of all, the scaling laws in AI have just held. A friend of mine pointed this out and it's a great point. If you look at the original Kaplan scaling laws paper — it went from 10^-9 to 10 petaflop days — and then concatenate additional compute from there to GPT-4, assuming some algorithmic progress, the scaling laws have held probably over 15 OOMs. It’s a rough calculation so it’s maybe even more. They’ve held for a lot of OOMs. They held for the specific loss function they're trained on, which is training the next token. Whereas the progress you are forecasting, we specifically know that that scaling can’t work because of the data wall. There's some new thing that has to happen, and I'm not sure whether you can extrapolate that same scaling curve to tell us whether these hobblings will also be fixed. Is this not on the same graph?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.