Evidence receipt / observation
Published · transcript-backedBronson Schoen: observation
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…So obviously, a ton going on there. Let me just do a couple quick calibration things. One, how much of the chain of thought that you've read has come from OpenAI specifically? I know you've done deeper collaborations with them and put out some papers. Is that the large bulk of what you've read, or how have like, how diverse is your chain of thought intake? Yeah. I would say Is there anything we should know about how they maybe compare as of now? I would say the the large bulk is OpenAI. We work with other labs. Can't say much about what what access we do. We don't have with respect to that. I think one good thing is that you're starting to see more research into the open models. One of the papers that I go back to a lot is called Metacot. It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great. We have credentials. We'll share credentials or whatever. And these are all valid. It's just that they take place in these just colossal explorations of different possibilities. And one of the other kind of pain points with this that seems to be true across models is that because the models are doing this broad exploration, it's difficult for any individual piece to be like, ah, they have concluded this. If the model's like, maybe we could do this, maybe something else, and it does 18 more considerations, then it's like, thus, the answer is a. And you're like, ah. And someone's like, why did it pick a? It's who knows, man? There's a bunch of kind of reasons. There's a similar paper called the thought branches where they basically look like resampling different parts of these sentences. And you can see that different kind of sentences depending on how you do the ablations have a different weighting effect on the final answer, but that it really does seem to be like this probability thing that's spread across a couple of different ideas, which I think intuitively makes sense. But as these chain of thoughts get longer and longer, this just becomes very difficult to boil down to a single reason. In the the recent UKAC Mythos preview incident, the, like, chains of thought for attempts on those evaluations are, like, about a 100,000,000 tokens apiece, which I think…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.