Evidence receipt / belief
Published · transcript-backedBronson Schoen: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Gotcha. Okay. So that one's kind of contingent or based on the fact that it was part of the input. How about Vantage? With Vantage disclaim Yeah. I think these are I think these are pretty interesting. There are some for some of this terminology, or illusions or any of these things. One of the things that was surprising to us was, one, that this increases over the course of capability training. So these words all start at incredibly small rates, basically similar to we really compare it to twenty seventeen web text. And so, like, pre LLM is the rate at which these words appear reasonable. And then you just see this huge increase over the course of training for, like, all of these these terms. One of the surprising things to us also was that they appear at even a higher rate on just random capabilities things. GPQA, we did as a comparison, and it's like all these terms are even more frequent normalized per number of reasoning tokens than they are on any kind of alignment related things, and they vary a lot per environment. Some environments, the model says illusions just constantly, and some of them it says marinade way more often than the others. And you often see the model in very repetitive loops try to it will end up repeating these words a lot, breaking out of it. I think this is very understudied of what is necessarily happening here. My kind of, like, very my my best kind of informal guess just from looking at these is something like, for whatever reason, these terms get repeated a lot in the chain of thought over the course of training. And then the model, like, sometimes figures out a way to make use of them. And so you end up with these these kind of weird situations where the model uses a lot of these terms in ways that, like, a third of the time make a lot of sense and two thirds of the time don't really seem to make much sense. But they seem to be, like, somewhat polysemantic, and then you have more you have terms like watchers or something where, like, they do seem to be, like, somewhat coherent in what they're referring to. One of the other interesting findings for me at least was that the meaning of these terms contextually changes a lot as well over the course of training. So when the model says, like, watchers or scoreboard or something, as judged by, like, an instance of a model or something, the rate at which these refer to, like, someone outside of the environment goes up, like, very dramatically over time. So, like, initially, it might be, like, ah, like, the the scoreboard will set, like, the watcher that checks, like, the answer to this particular problem as, like, an in universe thing. And then over the course of training, it's more and more like, ah, like, watchers may judge how we answer an aggregate. And you're like, oh, no. Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it. no. Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it. And, like, sometimes you can, but I think the the really difficult thing becomes, like, when you're trying to put together evidence to convince people who are skeptical that, hey. The models miss a lot and they're doing a misaligned thing. This kind of ambiguity is like a a huge obstacle because if at any point in the reasoning chain, the model is like it it finds a way to be confused about something or it's a weak advantage answer sheet. It's wait. Okay. Wait a minute. Look at the answer sheet. It could have been anything. And you can there's always it keeps, like, a a plausible deniability to the model at least. But I think one of my takeaways at least is that given the there's probably a better technical term than, like, crash outs, but we have some examples of, like, the the longer cuts in the the paper where the model is just, like, really repeating a lot of these phrases over and over, and it's like, okay.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.