High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Trenton Bricken: belief

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I think, again, if the model has the right context and scaffolding, it's starting to be able to do some really interesting things.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
belief
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…Right. Especially since your compute bottleneck anyway, so trying out his ideas… Or I guess he doesn't have the compute to try out all of his ideas. But Dwarkesh, you said, "Oh, well, the model can do the more straightforward things and not the deep thought." I do want to push back on that a little bit. I think, again, if the model has the right context and scaffolding, it's starting to be able to do some really interesting things. The Interp agent has been a surprise to people, even internally, at how good it is at finding the needle in the haystack when it plays the auditing game, finding this reward model bias feature, and then reasoning about it, and then systematically testing its hypotheses. So it looks at that feature, then it looks at similar features, it finds one with a preference for chocolate. It's like, "Huh, that's really weird that the model wants to add chocolate to recipes. Let me test it." So then it will make up like, "Hey, I'm trying to make a tomato soup. What would be a good ingredient for it?" And then sees that the model replies chocolate, reasons through it, and then keeps going, right? There is conceptual understanding. Deep conceptual understanding…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence