Evidence receipt / belief
Published · transcript-backedSholto Douglas: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“I can't really talk about where exactly people sit on that scaffold. I think different people, different tasks are on different points there.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Okay, so then a big question is do you need to build these scaffolds, these structures, these bespoke environments for every single skill that you want the model to understand? Then it's going to be a decade of grinding through these sub-skills? Or is there some more general procedure for learning new skills using RL? It's an efficiency question there. Obviously, if you could give a dense reward for every token, if you had a supervised example, then that's one of the best things you could have. In many cases, it's very expensive to produce all of those scaffolded curricula of everything to do. Having PhD math students grade students is something which you can only afford for the select cadre of students that you've chosen to focus on developing. You couldn't do that for all the language models in the world. First step is that obviously, that would be better. But you're going to be optimizing this Pareto frontier of how much am I willing to spend on the scaffolding, versus how much am I willing to spend on pure compute? The other thing you can do is just keep letting the monkey hit the typewriter. If you have a good enough end reward, then eventually, it will find its way. I can't really talk about where exactly people sit on that scaffold. I think different people, different tasks are on different points there. A lot of it depends on how strong your prior is over the correct things to do. But that's the equation you're optimizing. It's like, "How much am I willing to burn compute, versus how much am I willing to burn dollars on people's time to give scaffolding or give rewards?" Interesting. You say we're not willing to do this for LLMs, but we are for people. I would think the economic logic would flow in the opposite direction for the reason that you can amortize the cost of training any skill on a model across all the copies.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.