High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Sholto Douglas: belief

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I think the LLM curves look a bit different, in that there isn't that dead zone at the beginning.”

— Sholto Douglas

Source trail

Everything needed to verify it.

Speaker
Sholto Douglas
Attribution
Verified speaker
Claim type
belief
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…The chess analogy is interesting. Sorry, were you about to say something? Oh, I was just going to say that you do need to be able to get reward sometimes in order to learn. That's the complexity in some respects. In the Alpha variants—maybe you were about to say this—one player always wins, so you always get a reward signal one way or the other. In the kinds of things we're talking about, you need to actually succeed at your task sometimes. Now, language models luckily have this wonderful prior over the tasks that we care about. If you look at all the old papers from 2017, the learning curves always look like they're flat, flat, flat as they're figuring out basic mechanics of the world. Then there's this spike up as they learn to exploit the easy rewards. Then it's almost like a sigmoid in some respects. Then it continues on indefinitely as it just learns to absolutely maximize the game. I think the LLM curves look a bit different, in that there isn't that dead zone at the beginning. They already know how to solve some of the basic tasks. You get this initial spike. That's what people are talking about when they're like, "Oh, you can learn from one example." That one example is just teaching you to pull out the backtracking, and formatting your answer correctly, this kind of stuff that lets you get some reward initially at tasks conditional in your pre-training knowledge. The rest is probably you learning normal stuff. That's really interesting. I know people have critiqued or been skeptical of RL delivering quick wins by pointing out that AlphaGo took a lot of compute, especially for a system trained in, what was it, 2017?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence