High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Dwarkesh Patel: commitment

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“If this scale of compute increase can’t continue beyond 2030—not just because of chips, but also because of power and raw GDP even—then because we don't think we will get it by 2030 or 2028, then the probability per year just goes down a bunch.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
commitment
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…If you think the tokens are equivalent. You still get pretty substantial numbers, even with your 100 million H100s and you multiply that by 100, you're starting to get to pretty substantial numbers. This does mean that those models themselves will be somewhat compute bottlenecked in many respects. But these are relatively short-term changes in timelines of progress, basically. Yes, it's highly likely we get dramatically inference bottlenecked in 2027 and 2028. The impulse to that will then be “okay, they'll just try and churn out as many possible semiconductors as we can.” There'll be some lag there. A big part of how fast we can do that will depend on how much people are feeling the AGI in the next two years as they're building out fab capacity. A lot will depend on the Taiwan situation. Is Taiwan still producing all the fabs and chips? There's another dynamic which was a reason Ege and Tamay, when they were on the podcast, said that they were pessimistic. One, they think we're further away from solving these problems with long-context, coherent agency, advanced multimodality than you think. Their point is that the progress that's happened in the past over reasoning or something has required many orders of magnitude increase in compute. If this scale of compute increase can’t continue beyond 2030—not just because of chips, but also because of power and raw GDP even—then because we don't think we will get it by 2030 or 2028, then the probability per year just goes down a bunch. Yeah. This is like a bimodal distribution. A conversation that I had with Leopold turned into a section in Situational Awareness called “this decade or bust”, which is on exactly this topic. Basically for the next couple of years, we can dramatically increase our training compute. And RL is going to be so exciting this year because we can dramatically increase the amount of compute that we apply to it. This is also one of the reasons why the gap between say DeepSeek and o1 was so close at the beginning of the year because they were able to apply the same amount of compute to the RL process. That compute differential actually will be magnified over the course of the year.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence