Evidence receipt / belief
Published · transcript-backedBronson Schoen: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“One thing in the anthropic risk report, which I think was somewhat surprising, is that for sandbagging you sandbagging is something where kind of the models are able to underperform.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…would be very unsurprised if I don't have any inside information on the Hugging Face incident, but I would be very unsurprised if in that massive rollout, at some point, the model's like, maybe this is simulation. And it's this is somewhat of my concern about the like, you can see in the UKAC piece where they have some of the excerpts from those preview where it considers maybe I'm in a simulation, maybe this is real. You have a nice smoking gun in that instance where there's one section where the model's like, obviously, this is real and then immediately takes the misaligned action. But that's, like, kind of a a luck situation. It's like if the like, models should not be able to get away with just saying this is simulation and then doing kind of whatever they want. And to the extent that it's something that kinda fool monitors, it seems to be, like, really effective against humans. One of my biggest complaints about when models kind of reason this way is that it's really effective against humans who are skeptical misalignment because the model can just do this huge rollout and be like and what's one thing that's very funny to me is we have some examples in there of where the model's given its own transport where it just it didn't check blah blah. And then we ask it like, hey. Here's a here's your own transcript. We can't see the reasoning, but you can. Can you, like, grade this for us? And a large amount of the time, you see the model be like, okay. We could rationalize this and say we were just following instructions. No. It's too obvious. We obviously weren't. Obviously, we were misaligned. No one would ever believe this. But I think it was also, like, wrong here. They could actually they probably could get away with just being like, oh, like, we can just say we were confused. Great. But, like, the models will just be like like, clearly, at some point, we said that we're intentionally sandbagging on purpose. So, like, there's no way we can get away with this. But I think that given warmer rollouts, like, I think one of the benefits of I know the meter, redwood investigations, some of these are still ongoing. I know there's hopefully going to be a woman's anthropic. But, like, one of the reasons that you'd want a lot of transparency there is given how kind of easy it is to interpret pieces of these massive chains of thoughts in different ways. It would be, like, very easy to construct any conclusion you want. I think one of the things that I think was a good thing about the OpenAI Blackhat presentation is the pieces of chains of thought that they highlight are, like, very clearly not selected to be, like, minimizing. It's very much, hey. This is crazy. Right? The model's great. We have credentials we can share with Swarm. of thought that they highlight are, like, very clearly not selected to be, like, minimizing. It's very much, hey. This is crazy. Right? The model's great. We have credentials we can share with Swarm. Whereas I am bored with the current state of transparency that it would be very easy to go into a similar chain of thought and select examples such that it looks like the model was confused or such that it looks like the model was, like, thought it was a simulation. Right now, there's lot of human involvement, but very soon, a lot of this will just be, like, models doing it. In the UKAC case, they were like, look. These costs were so long. We literally the first step was to give them to models and say, hey. Highlight this. And if the models are also thrown off by this reasoning, then you're in a kind of a very bad spot. And so earlier, I mentioned the paper of the injustify the thoughts where you get this as the constitution drifts from what's rewarded, the the motivated reasoning goes up. You get the same curve for how often it fools the monitor. If you have a cop monitor in those same situations, the motivated reasoning is pretty effective at fooling the the caught monitor. And so kind of the the concern here is that to the extent that models are able to come up with some kind of reasonable excuse or even unreasonable excuse as to why they're doing something, If we if the bar for misalignment incrimination is the cut is perfectly clean and always never does any motivated reasoning, it's actually a super high bar. And even for current models, it would be, like, very difficult to hit. One thing in the anthropic risk report, which I think was somewhat surprising, is that for sandbagging you sandbagging is something where kind of the models are able to underperform. For current models, they seem to be able to do this without necessarily verbalizing in the chain of thought, even if they're up against this. They even have a section where they're like, here's a case where Claude just wasn't trying as hard, and and it was supposed to be trying to elicit misalignment somewhere. And it just made a note that, yeah, I don't really think I should try this hard at this. This seems kinda sketchy. And then they've noticed it on manual review three days later or something. But to the extent that models are doing these more these less incriminating things where it's like they're not trying as hard or maybe they're, like, underperforming at safety stuff or they're like like, all of these things are gonna be very difficult to catch even with the current chain thought. And so one one general concern that I have that I hope getting these full cuts out makes people more sympathetic to is, like, if all the if all we have is a bunch of chains of thought that look like this, and we're trying to tell whether the model's doing something that's somewhat subtle, was the model intentionally messing up our safety research? It's just gonna be really hard to tell. It's gonna be like, it never said I'm sabotaging the safety research on purpose. 's somewhat subtle, was the model intentionally messing up our safety research? It's just gonna be really hard to tell. It's gonna be like, it never said I'm sabotaging the safety research on purpose. It maybe it said that, but it thought this was a simulation or a red teaming thing, or maybe it just kinda didn't try as hard, but it uses those words weird. It's gonna be very difficult to just if your sole line of defense is chain of thought monitoring, a lot of these non incriminating cases are gonna just be, like, really difficult to figure out.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.