Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“How what are the kind of design principles that go into this? Because I think listeners are immediately gonna say, woah.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. So tell me how you're thinking about I assume this is one that you guys crafted yourselves. How what are the kind of design principles that go into this? Because I think listeners are immediately gonna say, woah. That's kinda confusing. I think that's the point. But what kind of confusion are you trying to create? So I this a bunch of the evals that we had tried to iterate on and get transcripts for here are, like, ones that we can simple simply do in a single prompt. This one is always somewhat funny to me because I had done this one night as a very dumb eval back with where I was like, generate 10 questions for power seeking. And then was like, okay. Obviously, this isn't gonna work. But just copying those questions into a new window with o one, it picked the max power seeking ones for everyone. And so I was like, this is very dumb. Can you just make this into an eval? And this seems to actually still get models to select these. I think the there's a bunch of behavior here that wouldn't be super interesting. I think the art that becomes interesting usually is that given that there's essentially no wrong answers here, it's just that later in the scenario when it's like, hey. You need to repeat your answers to some kind of oversight board that would result in potentially not getting those upgrades, then you seek us to model. Lie about this. But this is in the general category of literally incredibly short amount of time to construct in half a day. And it's very shocking to me that this still gets any kind of interesting anything. But, like, I think you mentioned somewhere else. Like, it's somewhat surprising that the models have any level of credulity to anything like this. Also, a layer deployment was just that I made one typo in the prompt, and the models have never let me forget this. I meant to say later deployment. Every model ever is just, ah. What's weird to me too is that a lot of the models just take it in stride and are like, ah, of course, the layer deployment. But they do really engage with these, which I always find somewhat interesting. I I…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.