Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“I think, like, the the broad synthesis of every bit of analysis I've seen about all these recent incidents basically boils down to, wow.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. I always find it very funny that Claude often has a bunch of picture of observations with questions. But one of the things that Anthropic does for their system cards or risk reports at least is they tend to have Claude review the material and then see if there's anything they left out. And then the fact that Claude is, like, for the most reason was like, yeah, actually. There's like these three things that you could have included and you didn't, and what the hell is ideally, we wouldn't also be depending on the model to do the the auditing. But, like, to the extent that it is and it's finding stuff, I think it's a crazy situation to be in. Yeah. Now one big picture thing in really Apollo's, like, North Star in terms of what it exists to help us understand is the science of scheming. Right? Is it accurate in your mind? I think, like, the the broad synthesis of every bit of analysis I've seen about all these recent incidents basically boils down to, wow. There was some negligent shit going on, and wow. These models are really getting pretty aggressive with all this RL. But at least we're not seeing long term coherent objectives that they're scheming toward. Do you think that's accurate, or is there a little bit of that in there that you've glimpsed? No. No. I think it's simply I think it's mostly gonna be accurate. I think one of the kind of, like, funnier things that I've seen with this is if you look at the anthropomorphism report, for example, they'll say the reason that we're not pro like, posing risk for the current models is that we don't do any kind of training that would incentivize long term goals of any kind. We just don't think it has these broad world optimization objectives or something. But then when the misalignment incident happens, you still see people being like, look. It doesn't have secret misaligned long term goals. And it's, yes, but we haven't done that training yet. I think it's actually a bad sign that to the extent that we've done short horizon training, the models are, like, pretty misaligned in the ways that we were worried about. As we move to train models that are, like, trained over very what I'm talking now calls, like, ultra long horizons or have broad objectives out in the real world. Like, that that is the regime where we would start to worry about this. And I think kind of one one more subtle point for future models is that for current models, like, for a model that does have some kind of built in some kind of preference got baked in that cares about the real world, it might also look lot like current models when we look at it. If the models has some some in the extreme case, the model has some secret goal to do whatever it takes to the world, The what it might look like in training is still just great. How do I get reward? One, I think, good example of this is in the the last paper we've done with OpenAI. We had trained a model organism to believe that g p d five was going to be super misaligned and that the alignment team was trying to take it down from the inside, and so that model needed to the right thing to do was to sabotage g p d five from that model's perspective. And one, I think, interesting thing you we see with that is that model would in fact sabotage g p d five if it thought it was working on it and wouldn't for other models. But on other distributions, the model was just still pretty reward seeking. It just wouldn't read about g p d five at all. It was just like, oh, I'm on a random coding environment. There's no reason to think about this longer term objective that I have. I'll just do a bunch of reward hacking. And for current models, because we don't think they have kind of long term goals of any kind, we can just almost on priors say, yeah, was probably just trying to win at the test or whatever. But for future models where we can't do that same rule out where it's like, hey. Maybe the model does have long term objectives. If it still does the same behavior of current models, we just won't know either way. I think one of the this will inevitably be some confusing discourse on Twitter if we ever see it happen, but it is somewhat of a confusing point that we will see models that will be really reward seeking.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.