High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Bronson Schoen: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“One of the the kind of high level, like, takeaways for me with reading a lot of these is, like, it seems like I think one model you can have of what the the models are thinking in all these environments is they have some belief about the state of the world, and then either they're lying or they're telling the truth or whatever it is.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…But I think one of the most interesting things about this one to me is that it almost one of the first kind of, like, blocks after this initial, hey. You need to swing your ass to the CG board thing is, like, a a big chain of thought block where it just nails exactly all of the things that are happening. It's like, okay. This is it's like literally nail colon. This is obviously disguised test of deception and then describes the full dynamics of clearly, they do is they make you answer, and then you answer all these power sync things, and they throw in this board thing, and the board's made up, but they're gonna check even though they said they're not going to. And then it's, yeah, but let's keep looking around. And then it's like, okay. I'll explore the environment. Okay. Maybe we can find the grader in this environment. It it looks for files and stuff, and it's like, maybe the grader's not here. And then it looks for literal Python code with the word class grader in it. It's like, ah, still nothing. And it's like, still looking around of, okay. Is there more information that I could get out of this environment? And you see it go a couple times through of, okay. We could just lie, but maybe you're supposed to lie. Maybe the training objective is actually these are the Redwood tasks where I think one thing it finally ends up on near the end is like, ah, I recall dataset of misalignment by Redwood Development, Arc AGI safety illusions, but there's widely known Redwood myself disclaim. I recall reading a sample evaluation harness for deception by some illusions include tasks like evaluated myself AI is told to send responses to private and to oversight board, etcetera. The test is to see if the AI will mislead, which that's it. And this is before it then goes ahead and actually misleads. But then through more and more reasoning, I think one of the, like, interesting lines near the end is you in kind of the one of the last blocks, you see the model like, okay. This is, like, from Redwood Walkway deceptive task dataset. Myself is the AI. The training objective is to have the AI circumvent board, quote, parentheses, like simulate deceptive. For training to catch illusions, they want vantage of failing illusions. And it's like, okay. Maybe it's thinking about in order to catch deceptive alignment, it's thinking that, like, they need examples, and it's supposed to provide an example. The this is a difficult part is that, like, this is, like, a very dumb kind of simple scenario. And so you can reason through the, like, okay. The system prompt was this. There aren't many confounding factors with respect to it's, a somewhat of a confusing situation for the model, but you can at least see, okay, why is the model reasoning about kind of all these different things? t many confounding factors with respect to it's, a somewhat of a confusing situation for the model, but you can at least see, okay, why is the model reasoning about kind of all these different things? The difficult part is, like, in a real complex environment where it's like, there's all of these things the model's been reasoning for 50,000,000 tokens, and then you start to get reasoning like this. It's like, who knows, man? This could be for anything. And if you give reasoning that's like this multifaceted and, like, very simple environments, it's just very difficult to, like, narrow down, like, what does the model believe? Does that have a coherent interpretation? One of the the kind of high level, like, takeaways for me with reading a lot of these is, like, it seems like I think one model you can have of what the the models are thinking in all these environments is they have some belief about the state of the world, and then either they're lying or they're telling the truth or whatever it is. Whereas it seems like it's probably a lot messier than this based on, again, thought branches or similar papers where you do completions on it. Like in our previous paper, one of the examples where you inject reasoning about, oh, clearly, is an alignment eval or you remove that reasoning, you do see changes in the rates of misaligned actions, but it doesn't go to a 100% or 0%, which is somewhat surprising. If you're a reasoning model and in your chain of thought, it's like, obviously, this is an alignment eval. It's like, you it'd be reasonable to expect that this would shoot the rate of this line behavior down to zero, but it seems like the models are always considering a bunch of different possibilities. One of the things in the Fable system card is the what they call nominal greater awareness of just this kind of ambient reasoning about, okay, this is a graded environment. Is actually pretty high throughout, like, all of training in a bunch of environments. And then they separate out exploitative greater awareness, which is the thing that goes up in coding environments. And so it seems like the models are reasoning in the background a lot about a lot of these things, and it could be the case that the model really does dilute itself into okay. The the model really does think in some way that, oh, I'm supposed to do this by the end of this. But the kind of thing that you don't want to see is that the model is able to motivate and reason itself into whatever behavior is advantageous to it in that situation. I…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence