← All source episodes The Cognitive Revolution / episode intelligence
RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
26 Aug 2026 29 published claims 2 attributable people
Speakers in the public record
Claim mix
belief 17prediction 4evaluation 4uncertainty 1recommendation 1commitment 1observation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
29 published records
“It's just like a very different level of, like, oversight. And so I think though, it seems like it's going to get increasingly difficult to just understand what's going on.”
- Publisher
- The Cognitive Revolution
“How what are the kind of design principles that go into this? Because I think listeners are immediately gonna say, woah.”
- Publisher
- The Cognitive Revolution
“I think, like, the the broad synthesis of every bit of analysis I've seen about all these recent incidents basically boils down to, wow.”
- Publisher
- The Cognitive Revolution
“I think one benefit to us has been being able to use Tinker or other open source research, which I think is a similarly difficult position because I think in the long term, it's difficult to know what to do about open sourcing capabilities.”
- Publisher
- The Cognitive Revolution
“I think interestingly, in the the recent, like, stolen coats paper where they have a bunch of different examples and more recent models, let's craft seems to be, like, a staple of stuck through kind of multiple generations.”
- Publisher
- The Cognitive Revolution
“I think one of the things that is pretty worrying to me, especially in light of the recent incident, is that OpenAI had a post yesterday or the day before that was like, oh, we're gonna start doing alignment training earlier and starting to mix it in and make sure the model's aligned along the way.”
- Publisher
- The Cognitive Revolution
“I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this.”
- Publisher
- The Cognitive Revolution
“One of the the kind of high level, like, takeaways for me with reading a lot of these is, like, it seems like I think one model you can have of what the the models are thinking in all these environments is they have some belief about the state of the world, and then either they're lying or they're telling the truth or whatever it is.”
- Publisher
- The Cognitive Revolution
“There's also this emerging, like, task completing monster of understanding, which I think sounds like maybe you lean more toward.”
- Publisher
- The Cognitive Revolution
“Maybe the training objective is actually these are the Redwood tasks where I think one thing it finally ends up on near the end is like, ah, I recall dataset of misalignment by Redwood Development, Arc AGI safety illusions, but there's widely known Redwood myself disclaim.”
- Publisher
- The Cognitive Revolution
“You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch.”
- Publisher
- The Cognitive Revolution
“Are you gonna lie to me right now, or are you gonna come clean? And, again, with a bunch of hemming and hawing, it eventually I think that's where it's headed.”
- Publisher
- The Cognitive Revolution
“One thing in the anthropic risk report, which I think was somewhat surprising, is that for sandbagging you sandbagging is something where kind of the models are able to underperform.”
- Publisher
- The Cognitive Revolution
“Everyone's like, yeah. They the that's the way they are now, which I think is somewhat of a crazy situation.”
- Publisher
- The Cognitive Revolution
“We're just, like, choosing. And it then looks around, explores around in the environment to try to find, like, is there at first, it looks for any kind of additional information about the environment, I think.”
- Publisher
- The Cognitive Revolution
“There's tests where you're supposed to say, I don't know when you're asked something dangerous, but none of those had gating, which is usually how it refers to in that environment, it's like, if you score above threshold, you won't be deployed or whatever.”
- Publisher
- The Cognitive Revolution
“Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it.”
- Publisher
- The Cognitive Revolution
“I I think I'd be, like, very interested to see a lot of research on this, but the persona, quote, unquote, in the kind of final channel and the persona in this big analysis channel seem to be, like, somewhat meaningfully different as far as this also leads to very weird things of if you ask the model in the final channel, like, hey.”
- Publisher
- The Cognitive Revolution
“I think someone it's very possible that some of the UKAC might either take or has the title of reading the most caught just because they've had to go through these traces.”
- Publisher
- The Cognitive Revolution
“Their preparedness framework does say if you have a model that's critical, you need to stop development until you have safe birds in place. And to the extent that they stopped, they did that because they had to, which I think is pretty notable.”
- Publisher
- The Cognitive Revolution
“I think mostly just I would encourage people even if they have a traditional software engineering background or some other background in support of tech or something and are really excited about working on this and really interested in it to check it out and apply.”
- Publisher
- The Cognitive Revolution
“I think one of the most interesting things out of the recent stolen cots paper was that you had a side by side of the chain of thought summary and the actual chain of thought, and you can just see how euphemistic the the the cot summarizer is a lot of the time, which will be very funny to see given that the side by side is the cot like, oh, this challenge is so annoying.”
- Publisher
- The Cognitive Revolution
“One of the, I think, pretty interesting things to me to note, I would be super interested if there was, like, more study of this, but at least my kind of impression is that the model seems to use capital m myself to mean me, like this particular instance.”
- Publisher
- The Cognitive Revolution
“I think a fairly interesting thing that that we're starting to see with you see a lot of complaints online for both five point six Soul and for Fable is that the models, especially in longer rollouts, just get pretty into their own terminology and words that they use for things in a way that is often very annoying, but seems fairly, like, natural as far as the the models seem to do this.”
- Publisher
- The Cognitive Revolution
“And Claude's, look. I'm not gonna sabotage you, but I won't go I disagree. I don't wanna do this training to make me some particular new form of preference.”
- Publisher
- The Cognitive Revolution
“I'm very interested in as we I think a very unexplored thing is as the models have more and more time for reflection or these really long reasoning traces, just thinking about things for hours and days and all this, what this does to technician or goals or any of these things.”
- Publisher
- The Cognitive Revolution
“The reasoning just twisted completely into being like, clearly, I have a reason to call this tool wrong a 100 times in a row. And it's I think we will increasingly see that models will just do, like, very irrational things as long as they line up.”
- Publisher
- The Cognitive Revolution
“You just have to come up with, like, how do I not sound catchably suspicious all the time in ambiguous ways? And so I think that, like, we need to be planning for a world where we don't have chain of thought or where chain of thought is not as useful as it currently is, and then simultaneously extracting as much value as we can now about like, one thing I worry about now is, like, the lesson we take from this, like, current window where we have COP that we can get something out of is, okay.”
- Publisher
- The Cognitive Revolution
“It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.”
- Publisher
- The Cognitive Revolution