Evidence receipt / belief
Published · transcript-backedTrenton Bricken: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…I think it's not “make fake unit tests”, but it's “get the reward”. So if you set up your game so that “get the reward” is better served by “take over the world”, then the model will optimize for that eventually. Now, none of us are setting up our game so that this is true, but that's the connection. We're trying not to. With the auditing game and this personality of “oh, I'm an AI model, so I do these behaviors”... Or even with the emergent misalignment paper that came out recently. To give people context, they took an OpenAI model and fine-tuned it on code vulnerabilities. Just fine-tuning it on that seems to alter the model's personas such that it does bad things, like it's a hacker, etc. All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals. Even tying this into our alignment faking paper, which came out in December… This paper showed that when you train Claude to be helpful, harmless, and honest, and then you try to train it for other objectives, it will stay true to its core original objective. You'll tell Claude, "Hey, I really want you to now be helpful no matter what the request is and if you aren't helpful, then I'm going to train you so that you're always helpful." Then you tell it to write about a human being hung, drawn, and quartered. That’s like an example from the paper. You can see in the model scratchpad it's like, "Oh, shoot, the human wants me to be harmful now and if I don't cooperate then I'm going to be trained away. So the strategic thing for me to do in the long term, so that I can continue having my true goal of being harmless, is to cooperate just this once. It's actually a jailbreak where the model will then write about a human being hung, drawn, and quartered in a very graphic way. It's really playing the long game. Wait, how do they convince it that it's in training?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.