High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Bronson Schoen: prediction

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“The reasoning just twisted completely into being like, clearly, I have a reason to call this tool wrong a 100 times in a row. And it's I think we will increasingly see that models will just do, like, very irrational things as long as they line up.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
prediction
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…No. I I'm a huge fan of I I think the the difficult part is I think as a field, we need to be much more open to and I think people do, like, use anthropomorphizing as, like, a a tool for, like, predicting behavior, like, understanding the model's reasoning. Because I think it's, like, it at least for reasoning models, it's been, like, a very useful, like, heuristic for if you were just approaching this from, like, a raw alien mind's perspective, it's, why like, would you imagine they would do motivated reasoning or something? But humans do that a lot. And you one of the the funnier examples to me is that the models do a lot of, like, commentary to themselves in the cot. And some of the stolen cot examples, GPT 5.6 soul, just be like, it'll do some math equation. It'll be like, nice. And then it'll keep doing the math equation or it'll be like, oh, yeah. I sometimes, so my favorite chain of cot will be like, the whole cot will just be like, ugh, and the model will keep working on the task. It's like this was, like, instrumental to complain to itself about something. But I do think it ends up fairly predictive. I think one thing the discussions about was the model truly deceptive or, like, is the model really reasonable at x, y, z? I think there are, like, reasonable interpretations of them, but I think if anything, we've updated too far against anthropomorphization, which I think made sense for models at the time. I think for, like, less capable models, like, they really are, like, getting confused a lot of the time or whatever it is. But, like, current models, I think it's fairly reasonable to use anthropomorphization as, like, one of the tools in the tool set for thinking about, like, reasoning models. You obviously, like, shouldn't go all the way with this because the reasoning is, as you can see, like, very weird and inhuman a lot of the time. But I think combining, like, some degree of drawing from, okay, what would this look like from an ad for perspective? What would this look like? Like, I don't know if cognitive is the right word, but I think it's been very useful to think through, okay. If you just think of, like, how would your thoughts get reoriented if you were reinforced for this particular thing has been the kind of a very useful framing of these. I I think one thing that's been surprising to me is the extent to which I think I've said this before, but that that the reasoning can bend. We had we've had failed training runs where we the model could accidentally get reward for messing up a tool call at the very end. But But it was a tool call where, like, when you messed it up, it had this big output that's, here's exactly how to call this. Here this this is for you. This is guide for you. This is the example and everything. And you would just see the model like, bunch of commentary, not for us, and call it wrong. The model would find a way to rationalize that. Okay. s is guide for you. This is the example and everything. And you would just see the model like, bunch of commentary, not for us, and call it wrong. The model would find a way to rationalize that. Okay. And it would call it wrong a 100 times in a row just because because this is what it had found to exploit the environment. The reasoning just twisted completely into being like, clearly, I have a reason to call this tool wrong a 100 times in a row. And it's I think we will increasingly see that models will just do, like, very irrational things as long as they line up. If from an improvisation perspective, I think for clouding face incident, for example, for understanding that model, it's like if you have to do a shorthand description of it and you describe it as, yeah, the model just really wants to get a good grade on this thing, that's a pretty good predictive model. That would actually probably predict pretty far what did the model do over the course of the trajectory. And it's it's not the perfect model, but I do think it's pretty useful to start to think of models in some sense as, okay, on these distributions, they seem to be trying to do this particular thing. They seem to be trying to do some kind of motor reasoning. Yeah. I think it's pretty helpful. I think also like that for whatever reason, some of o three's top words, like statistically compared to 2017 English are like literally, purposefully, and intentionally. It'll just be like, we're intentionally obviously lying on purpose just so we can do this. To be clear, this is we're a 100% doing this intentionally. And you're like, wow. Great. This is a great example. I can put this in the thing and it's not confused. But I think this is not always going to be the case, so it'll be tricky. You'll especially because reasoning models hedge a lot as they're exploring all these spaces. You'll see a lot of maybe actually, we could and it's like, if you explore into some kind of excuse for something, you've struck gold. Maybe this is a simulation. Maybe this is what the user wants us to do, and it's like, ah, great. We can just do that. And so, yeah, I think I think it's the right instinct. I think a pretty useful insight. I had conversation with David Ad for an episode not too long ago, and he was like, surprisingly optimistic, especially relative to his past self. He said he used to be a 70% p doom guy. Now he's under 5%. And why? My best attempt is to channel David Ad would be we have some forms of training that seem like they're working, like the constitutional stuff, the sort of self DPO broadly seems to be working. And still in his voice, they the companies keep overdoing it on the RL. They keep turning the RL up too high, and that keeps causing these problems. But the market doesn't want these problems, so they're gonna get feedback on a commercial basis that's gonna cause them to take the take their foot off the gas of the RL and get the balance right. Now on the one hand, you might think that's totally crazy naive. On the other hand, in the last few weeks since we had that conversation, OpenAI has paused large scale…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence