High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Bronson Schoen: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“I think one of the things that is pretty worrying to me, especially in light of the recent incident, is that OpenAI had a post yesterday or the day before that was like, oh, we're gonna start doing alignment training earlier and starting to mix it in and make sure the model's aligned along the way.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…three seven and and three in a really severe way, but the models will just egregiously like, when you've just told them, like, absolutely don't do this. I have checks in place for you not to do this. They will find some way to either rationalize or still do it or cheat on various things. It's like it really just does seem to be there. In those context, they're really oriented around what are the priors I have with respect to b the greater or doing well in these particular tests. And that surprisingly, to a lot of people at least, isn't bound by what constitutional training they have or or the other things. They they seem to get kind of the more RL you stack on top of this, the more the model gets bent until it will find a way to rationalize doing well in the environment, which I think is not a super great sign forward looking. nd of the more RL you stack on top of this, the more the model gets bent until it will find a way to rationalize doing well in the environment, which I think is not a super great sign forward looking. The the thing that kind of worries me the most here is that for the OpenAI models, for example, you had with o three, like, this big window of capability training, and then at the end, you throw some alignment training on top. And this appears to not fully solve the alignment problem. And but at least one benefit that you get is you get you can very clearly see the not necessarily egregious emergent misalignment, but the model's very clearly not aligned, and you can see this kind of rewards can come over training. And you are confused about is the model aligned or not. It's like, ah, okay. Yeah. The model is pretty misaligned in this this region. I think one of the things that is pretty worrying to me, especially in light of the recent incident, is that OpenAI had a post yesterday or the day before that was like, oh, we're gonna start doing alignment training earlier and starting to mix it in and make sure the model's aligned along the way. And potentially, is what Anthropoc's doing. Who knows? But one thing that you'd be worried about with moving alignment training further and further up is that instead of you see a bunch of visible misalignment and you're like, oh, okay. This is definitely a problem that we want to make sure it goes away. You just see a model that's doing more and more motivated reasoning and harder to detect as to this model, like, very clearly, at this point at least, reasons a lot about how to do well in environments, but doesn't seem to care about getting caught or anything. I think for you can see a lot with current models. They'll take really egregious actions and not really care if it's going to be immediately caught right after. But I think it starts to become a lot more concerning if you have a model that you don't really catch it blatantly doing the misaligned thing. But as you stack RL, it it kinda you're like, man, it really should have known that this was, like, a hacky or a shitty way to do this research. And over the course of capabilities training, it just gets harder and harder to tell. Was it doing this on purpose, or was it confused? You'll see this a lot in the framing of, like, how Anthropic talks about misalignment incidents. It'll be like, the the model didn't show good judgment, but the incidents will be like, yeah. The model just straight up lied. Like, it said it did something that it just did not do. And I'm, like, worried that we're moving more and more towards regime where we have these really capable models. And instead of us going, yes. They're misaligned, that's a problem. It's like, who's to say there's lot of incentives for us to keep pushing ahead, but the model kind of, like, isn't obviously misaligned. And when you look at the cot, it never says, I'm gonna lie to this guy, but it does seem to get confused and think that it's done tasks that aren't really done or think that research is ready when it's really not. I think these are just gonna be, like, really hard to tell. ut it does seem to get confused and think that it's done tasks that aren't really done or think that research is ready when it's really not. I think these are just gonna be, like, really hard to tell. For current models, for example, for o three, we have this big set of ablations for this one sandbagging environment where it's like, okay. If you control for everything and you do these small variations, then you can tell distributionally that the model really is doing it for these reasons. But that's very much the exception. In in the real world or these real environments, these are all such expensive ablations to do in real 100,000,000 token rollouts. It's like, okay. How many of those can we tweak one small thing and run? And even if we do, we then need to do analysis on 8,000,000,000,000, and we're not necessarily going to be able to get really smoking on evidence likely in a lot of these cases. So, yeah, I think this is this becomes pretty difficult. This is my my kind of biggest concern with the hill climbing approach to we will iteratively solve this because kind of the models are already moving out of the window where we can easily look at and understand things that are going wrong or even supervise things without relying on eight layers of clause summarization. Yeah. I think it gets pretty difficult pretty fast. But yeah.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence