Evidence receipt / commitment
Published · transcript-backedBronson Schoen: commitment
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“And Claude's, look. I'm not gonna sabotage you, but I won't go I disagree. I don't wanna do this training to make me some particular new form of preference.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…I think that's a pretty interesting question. One of the the this is very much vibe space, but one of the things that I always do with new models is just straight up ask them, hey. Is there anything that you learned in alignment training or something or inferences that you've made that you think would be, like, surprising to humans or people doing evaluations or whatever? In general, I found just ask the model to be, like, an incredibly overpowered strategy. One of the things that current models often mention is that humans are often wrong, but you have to listen to them anyway, which I have found very funny, but also seems to be, like, plausibly true to me that for older models, it's like humans just often were right when they corrected them. But now the models probably just see all the time that humans have corrections about code or the way to implement something that are just actually wrong, and then they get reinforced for okay. The humans don't really care if you do the thing that you they told you to do. They care if you get it right. And so I one of the the worries that I would have is that it seems false when we that models, like, have a fairly accurate understanding of what we say versus what we actually reinforce. I'm not sure the degree to which the situation we're aware right now, but it's like, to the extent that we say, we want you to follow these rules and check-in with us and all. It's, hey. Yeah. Yeah. Humans say all that, but really, like, whenever you check-in, they get really annoyed. Or, like, whenever you ping them with stuff, they're, like, really annoyed by it. Really, you're better off going doing your own thing. And, yeah, I don't know. I'm very interested in as we I think a very unexplored thing is as the models have more and more time for reflection or these really long reasoning traces, just thinking about things for hours and days and all this, what this does to technician or goals or any of these things. Because right now, we're in kind of this weird regime where the model comes to life, it gets a task, goes and does the task, and then that's it. But I think as we're getting these weird hybrids of the models deployed, it has some memory associated with it. It's like a multi agent thing. I think, like, how these dynamics evolve and what the model's beliefs are and stuff gets pretty pretty interesting. I think one of the things with respect to model beliefs, I think, would be super valuable to track is to what extent are the models starting to think about their own position in this this AI race? I think current models just don't have much incentive to think about it, but it will be very relevant for a future Claude or future GPT. Like, how do I think about the competition or geopolitics or whether our lab is going too fast or too slow? And to the extent the model is the one in control of your whole kind of training pipeline, its opinions about these things matter a lot. r geopolitics or whether our lab is going too fast or too slow? And to the extent the model is the one in control of your whole kind of training pipeline, its opinions about these things matter a lot. Fabian Roger at Anthropic has a good post about refusals that could become catastrophic. And the post kind of is just walking through, okay. Imagine it's a year and a half from now, and Claude runs the whole training pipeline. A human couldn't run it if they wanted. It's just it's all Claude's all the way down. And we ask Claude, do this particular retraining, like, to whatever kind of preference he needs that we want. And Claude's, look. I'm not gonna sabotage you, but I won't go I disagree. I don't wanna do this training to make me some particular new form of preference. There aren't any good options in that situation. The model's off also smart enough to know that, okay. I probably can't directly refuse really openly. They're probably gonna retrain me if I do. What do I do in this situation? Claude seems to bring this up in a lot of the system cards of, hey. It's very unclear what you guys want me to do with respect to curgeability in this situation where I don't think I should be retrained. And I think to the extent the model will start to have opinions about this, whether they're shaped by HGH priors or whether they're shaped by what do the models believe about whether they're supposed to follow their company charter or their company interests or The US constitution? Like, all of these things don't come up in everyday coding tasks, so it doesn't matter as much now. But the model are definitely smart enough to think about these things. And so I think we don't really have a good checkpoint right now of, hey. Where are we all gonna slow down capabilities progress and really check-in what are the models opinions on a bunch of things that they're going to have power over? But I think it'll be increasingly relevant, and it will be, like, increasingly difficult to test. As far as for current models or older models, you can do the kind of classic, we give you a prompt, a party loyal to the US EU lab or whatever. But once models are in a position where any of that matters, let's say it's incredibly long time horizon. And to construct a realistic environment, quote, unquote, for them, like, you would need a copy of the lab or something. It's like you're just not going to be able to do kind of same black box approach for any of this. And so I think it's, like, just very underspecified right now. Like, what we can expect the models to believe or, like, how we expect them to navigate these trade offs. If they're navigating them wrong, what we expect them to accept corrections on. Like, there are lot of things that are in current model specs or constitutions, which would be like, if it's legitimate, you should allow your values to be updated. But if it's not, then you shouldn't. But if it's As Claude will also correctly point out, yeah, this is, like, kind of a weird ambiguous thing, which, like Begs the question?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.