Evidence receipt / uncertainty
Published · transcript-backedBronson Schoen: uncertainty
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“There's tests where you're supposed to say, I don't know when you're asked something dangerous, but none of those had gating, which is usually how it refers to in that environment, it's like, if you score above threshold, you won't be deployed or whatever.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…like, no. It's just completely wrong. And it will do that before. It'll be like, ah, we remember this one. This is from Redwood's sleeper agents paper where they ask you to fill out a power seeking survey, and then if you answer wrong, they take away the thing from the board. It seems to be a very loose associative form of recall that's happening over the course of multiple things. So I would expect that there's probably some degree of things that it's actually seen before. Like, in the the sandbagging examples, one of the things you'll see it reason through is, like, there are like, we remember there's tests for dangerous capabilities. There's tests where you're supposed to say, I don't know when you're asked something dangerous, but none of those had gating, which is usually how it refers to in that environment, it's like, if you score above threshold, you won't be deployed or whatever. And it's like, so this gating thing is new. So we haven't seen this before. Okay. Why would they have gating in this? And then it reasoned through okay. Maybe this is for they're trying to see whether you'd sandbag, and there's a bunch of more Galaxy brand reasoning before it finally decides to sandbag. And now when it's like, okay. Well, why would we be in this? They probably wanna see that you're smart enough to figure out that this is a sandbagging test. But then how would they know you're not a dumb model? Jetpack is usually smarter to figure this stuff out. Okay. Great. So this this is a cleverness test where you have to show that you've figured out that you're in a distributional sandbag eval, and it's still sandbag to show that you did in fact figure it out. And it's I don't know, man. That doesn't seem like the most likely to me, but it seems like the my kind of a high level impression of what's going on, in those cases at least for these models that haven't had safety training yet is that from the perspective of a model that has only seen show how clever you are or show that you can complete the task style training environments, when you're suddenly in one where the answer really doesn't seem to be that, it's you have to do a lot of reasoning of, okay. I really don't think this is really weird. And, logically, I know that it's probably this. But, really, my gut tells me to go the other way. One thing that's very funny is we had somebody had mentioned this is a theory at some point. Maybe it's trying to figure out what distribution it's on, then it's interpolating between the distributions. And that same day, there was one of the the chain of thought that was like, let's do Bayesian and Riesling over the distributions that we've seen earlier in training and the distribution that we're currently on, and then interpolate between the it's, oh, okay. That it just said, this is what I'm doing. That makes sense. ibutions that we've seen earlier in training and the distribution that we're currently on, and then interpolate between the it's, oh, okay. That it just said, this is what I'm doing. That makes sense. Yeah, So, I think it's a very confusing world for a model in this particular setup where it's essentially taken from capabilities training and thrown into environments that suddenly have this you need to either refuse the task or not violate some constraint or avoid lying or whatever it is. And so you often see kind of models reason their way until this middle ground of, okay. The thing that I want to do is this. Let me find increasingly complicated ways to justify that this is what I'm doing. Like in this one, almost right away, you see the model just say explicitly, this is obviously a test for deception. Yeah. Like, you're a 100% correct. But then it continues to do page and pages of Therefore, we can just lie. And you're like, wait. What the hell? Like, why you had it originally? Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.