High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Sholto Douglas: belief

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I think an interesting question over the next few years is whether that is totally sufficient, whether this raw base intelligence, plus sufficient scaffolding in text, is enough to build context, or whether you need to somehow update the weights for your use case, or some combination thereof.”

— Sholto Douglas

Source trail

Everything needed to verify it.

Speaker
Sholto Douglas
Attribution
Verified speaker
Claim type
belief
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…My question is always, are you giving the model enough context? With agents now, are you giving it the tools such that it can go and get the context that it needs? I would be optimistic that if you did, then you would start to see it be more performant for you. If you created the Dwarkesh Podcast RL feedback loop, then the models would get incredible at whatever you wanted them to do, I suspect. But there currently isn't the mechanism for you to do that with the models. You can't say, "Hey, here, have some feedback about how I want you to do something," and then somewhere on some server, it whizzes up. Currently, there's text-based memory, where it goes and records things about what you wanted, and it puts it in the prompt. It tries to build its own scaffolding in context. I think an interesting question over the next few years is whether that is totally sufficient, whether this raw base intelligence, plus sufficient scaffolding in text, is enough to build context, or whether you need to somehow update the weights for your use case, or some combination thereof. So far, we've only explored the first. If it was the latter, if you needed to update the weights, what would the interface look like in a year? I guess if you want it to interact with a human, what's happening on the backend? Is it writing practice problems for itself? Is it building actual environments for itself that it can train on?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence