Evidence receipt / belief
Published · transcript-backedSholto Douglas: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“In particular I think the intellectual ceiling goes quite—contra what I was saying before, which is we've demonstrated this incredible complexity of math, and programming problems… I do think that the type of task and setting that AlphaZero worked in this two-player perfect information game basically is incredibly friendly to RL algorithms.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, exactly. In retrospect, obviously a bunch of the methods are still used today in deep learning, and you can see similar things in the models that we train today. But it was fundamentally not a baby AGI that we just had to add a little sprinkle of something else on top of in order to make it the LLMs of today. I just want to very directly address this crux of, why are LLMs in a much different position of respect to true AGI than AlphaZero? Why are they actually the base on which adding in a few extra drops of this kind of care, and attention gets us to human-level intelligence? I think one important point is that when you look at AlphaZero, it does have all of those ingredients. In particular I think the intellectual ceiling goes quite—contra what I was saying before, which is we've demonstrated this incredible complexity of math, and programming problems… I do think that the type of task and setting that AlphaZero worked in this two-player perfect information game basically is incredibly friendly to RL algorithms. The reason it took so long to get to a more proto-AGI style models is you do need to crack that general conceptual understanding of the world, and language, and this kind of stuff, and you need to get the initial reward signal on tasks that you care about in the real world, which are harder to specify than games. I think then that sort of gradient signal that comes from the real world, all of a sudden you get access to it, and you can start climbing it, whereas Alpha Zero didn't ever have the first rung to pull on. Yeah, yeah. This goes back to the monkeys on the typewriter and the pre-training model. Until you had something like GPT-3/GPT-4, it just couldn't generate coherent enough sentences to even begin to do RLHF, and tell it what you liked and didn't like.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.