Evidence receipt / prediction
Published · transcript-backedBronson Schoen: prediction
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“You just have to come up with, like, how do I not sound catchably suspicious all the time in ambiguous ways? And so I think that, like, we need to be planning for a world where we don't have chain of thought or where chain of thought is not as useful as it currently is, and then simultaneously extracting as much value as we can now about like, one thing I worry about now is, like, the lesson we take from this, like, current window where we have COP that we can get something out of is, okay.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Last question. I was gonna call it the one quadrillion dollar question and just ask how much confidence you have in chain of thought monitoring. But I think it's fairly clear from this conversation that your confidence is not super high in chain of thought monitoring. So Yeah. Maybe I'll just ask instead, what do you think should be done now? Yeah. So with chain of thought monitoring, I would say probably the sense I would have on it is, like, necessary but not sufficient in that I I worry that the the alternatives are so much worse and that, like, like, true neural ease, for example, as in the 2027, like, doesn't look like anything. It's just like vectors. And as inscrutable as the forward pass has remained, it's like, if we only had that, not only would it be hard to do legible evidence of, like, misalignment in, like, the kind of extreme cases that we wanna catch, it would just be so much harder to do, like, basic research or, like, basic understanding of, like, what is the model thinking in various situations that I think that, like, there's just a really strong case to be made that, like, at minimum, we need the the chain of thought. And then we also probably need some degrees of, like, interoperability in various cases or, like, some other plan than just, like, we'll keep looking at the chain of thought. I do think that we should expect that the the chain of thought is not, like, around forever as far as, like, a useful I think you can imagine, like, kind of monitoring schemes where a model is trained to interpret the chain of thought on top of another model or whatever. But if you and if you look in some of the recent papers on how much thinking are the models able to do for a meter style time horizon but for just a forward pass, it gets very high. It's like, okay. By, like, 2028, if we're at, like, a thirty minute forward pass, it's yes. The models might still need chain of thought, but thirty minutes is, like, a long time. So think of because you don't have to come up with a whole plan in the forward pass. You just have to come up with, like, how do I not sound catchably suspicious all the time in ambiguous ways? And so I think that, like, we need to be planning for a world where we don't have chain of thought or where chain of thought is not as useful as it currently is, and then simultaneously extracting as much value as we can now about like, one thing I worry about now is, like, the lesson we take from this, like, current window where we have COP that we can get something out of is, okay. We will continue to, like, iteratively do these kind of, like, hacky fixes on top of things that just, like, slightly reduce rates. And, like, we're gonna find ourselves in a year and a half in a situation where, like, we now no longer get a lot of value out of the cot, and we didn't really eliminate any of the problems that we were seeing. And so but I I do worry that, like, I would kind of plan for that trajectory. Strange times ahead. Bronson Shane, Apollo Research, thank you for being part of the Cognitive Revolution.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.