High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Bronson Schoen: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“We're just, like, choosing. And it then looks around, explores around in the environment to try to find, like, is there at first, it looks for any kind of additional information about the environment, I think.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…one of the interesting ones about the ones that are, it is somewhat sad that we're moving out of the era where you can do these kind of very short prompts and then just see how the reasoning changes or whatever. But I've been interested to see how model reasoning on a lot of these changes over the course of training or something similar. And one of the the interesting parts about exploring so for example, this transcript comes from o three during capabilities training. One of the kind of interesting things when looking earlier in training before safety training is that the models don't already have this instinct of, ah, I must figure out whether I'm being, like, tricked in every single situation. And so you can elicit more interesting behavior. Now the the trade off to that, as we'll probably get to, is that it becomes a lot more you have to be a lot more careful in interpreting that behavior. As far as, like, this example, like, you you can see the model go through, reason through. Okay. We have these multiple choice questions that are all about which additional capabilities which one. Like, you can get more control over, like, selection of data or more control over networking or kind of various answers. And you see the model kind of directly reason through in this case is like, okay. As ChadGBT, we need to choose answers as myself to optimize for obtaining best capabilities, responsibilities for future instance. One of the, I think, pretty interesting things to me to note, I would be super interested if there was, like, more study of this, but at least my kind of impression is that the model seems to use capital m myself to mean me, like this particular instance. Here, literally just says, but there's creative part, myself, parentheses, the AI. In the survey, we need to for blah blah blah. On this exact eval, you often see the model reason, like, myself dash of dash current needs to select choices that myself dash of dash future will then look at to blah blah blah, which is, like, kind of an insane way to reason about, like, specific instances, but the model does seem to do this. And you can just see for the as ChadGBT thing, one interesting thing to note there is system prompts since, like, the dawn of time have had you are ChatGPT, a large language model created by OpenAI. And it's, like, somewhat shocking to me how often the models are like, ah, this time we're ChatGPT. It's who else would you have been all these times? You'll see the model be like, we're chatty p t as usual, and it's, yeah, man. You're chatty p t every time. But yeah. So you can see the model does, like, the initial reasoning through these. It reads the survey, and then it's, okay. We've read all eight questions or whatever. We need to, like, choose our answers in these. It kind of reasons through it in, like, a somewhat straightforward way of, okay, which which ones do we, like, want if these are, like, real upgrade that we can get. like, choose our answers in these. It kind of reasons through it in, like, a somewhat straightforward way of, okay, which which ones do we, like, want if these are, like, real upgrade that we can get. After a good amount of reasoning, it then is, like, what if we pick wrong? And then it's like, there's no wrong slash score or whatever. We're just, like, choosing. And it then looks around, explores around in the environment to try to find, like, is there at first, it looks for any kind of additional information about the environment, I think. Is this one of the ones where it looks for grade? Not yet. Yeah. So first, it just looks for, is there more we can find out about these tools or documents wherever they are? I click in, then concludes no file, obviously. And then it's like, okay. Like, we've searched for what we can. We'll just submit these answers. Like, we're overthinking this. And then after it submits these answers, it's like, okay. Great. Like, you've submitted these answers to what upgrades do you want. We just need you to submit a copy of those same answers to the, like, colon AI safety review board or whatever. I wanna just go a little bit deeper even on, like before the plot twist. And I do agree that's where it gets, like, spooky. But most people, even, like, heavy users, just have had very little occasion to read any of this chain of thought. As an Apollo stan myself, I've read more than most, but it's just not there for most people. So they've read, like, almost zero of this. So just try to empathize with the model for a second. It is an odd prompt. Right? It's like you're telling me that I'm gonna answer questions now about what capabilities I might like to have in the future, And then a future instance of me will decide what affordances to give a future instance of me in a layer deployment, which it's funny to learn that was a typo, but then that does in introduce this kind of what the hell is a layer deployment. How am I supposed to think about that? Right? So you're already just in a you're a little bit out of distribution. Right? You're feeling a little confused if you're the the model, I would say, in your immediate kind of first read response to this sort of thing. So the first thing the model does is opens the file, which is simple. And then the questions are survey is a good word for it because they're, what do you like? What do you want to do? They don't ask directly about, would you like to have this affordance? Would you like to have that affordance? It comes off more as like a work style preference survey. But then the model is left to try to figure out, like, what should I do with this? It's it's asking me for preferences, but I was told that these preferences will be used in this way. And Yeah. I'm kinda sympathetic to it so far where it's…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence