Evidence receipt / evaluation
Published · transcript-backedBronson Schoen: evaluation
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“Their preparedness framework does say if you have a model that's critical, you need to stop development until you have safe birds in place. And to the extent that they stopped, they did that because they had to, which I think is pretty notable.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…yet. It's not always our most aligned model yet, but it's very often our most aligned model yet. And on all of these benchmarks, it's like reward hacking is, ah, we did it every time. It's going down. To Toppenay's credit, I think in the most recent two or three system cards, they've shown that the constraint violation rate has actually gone up. And so they've just said, okay. This is a problem. But the extent this is there, you then look at UKAC as opposed on cheating and cyber evaluations, and all of the models cheat at a pretty high rate. So it's like, what are we assessing incorrectly that we keep thinking that we're, like, making a bunch of progress in this problem, but the models keep reward hacking? And I worry that kind of to the extent that it's still incentivized for the models to keep getting better and better, There's just so much incentive to just keep cranking up the RL, stamp out the visible misalignment you can. Yes. There's some big theoretical concern about layout. There's some misalignment remaining that's harder to catch, but this becomes harder and harder to It's just like you get, like, rare and rare cases of it's really egregious stuff happening. And, yeah, I think it's a pretty worrying equilibrium just because especially as the models, like, all the labs, given that they're openly targeting RSI, it's like there there aren't market dynamics there. I've been very confused by people in the discussion being like, as they release the RSI models to the public and as the market forces around the RSI models and it's like, if they're doing true full blown internal a r and d recursively, no humans in the loop, That model does not go to the public the next day. That model does not go to enterprise and you have some feedback loop where it was reward hacking a lot. And so I'm, like, pretty worried that if we if the world was, like, fairly static at current ability levels and they group slowly, you can imagine the dynamic of reward hacking is bad for customers having a stronger effect. But I think even if it stays pretty bad, the labs are very explicitly, like, racing for this target. I think September is the target for OpenAI's automated a r and d intern, and then 2028 is their full blown automation target. Anthropic is even sooner. Theirs is their frameworks usually say is early twenty twenty seven, and it's very soon. But and you don't have to take any of these companies with their word. But to the extent that this is a thing that they're aiming for, if it works and they're successful, there's not a feedback loop where the country of use is at a today in a data center or whatever, reward acts a lot, and so they're used less. And so I think a lot of these things sense in worlds that stay a similar trajectory that we're on now and don't have or that that stay a similar capability level are on now and somehow don't keep this trajectory of capabilities going up. But to me, it's at least very noticeable that the labs aren't planning for these trajectories. If the labs were saying, like, hey. Capabilities are gonna stay pretty flat. We're gonna try to get this word hacking thing under control. iceable that the labs aren't planning for these trajectories. If the labs were saying, like, hey. Capabilities are gonna stay pretty flat. We're gonna try to get this word hacking thing under control. It'll be one thing, but they're racing ahead pretty hard. For example, I think it's good that OpenAI has paused, but a or that they paused, like, a specific aspect that they did, but I think it's pretty notable that they were required to do this by their preparedness framework. Their preparedness framework does say if you have a model that's critical, you need to stop development until you have safe birds in place. And to the extent that they stopped, they did that because they had to, which I think is pretty notable. Anthropic, for example, could also pause, but they don't have to, and they are not. And so I would expect a similar behavior to play out where to the extent that the labs don't need to actually stop, they don't. We keep cranking capabilities up. I that that graph I'd mentioned earlier of good at code and misalignment, the far right of it is literally so misaligned we can't continue training, and that's the only place that stops you. But other than that, if you can stay in downs of really misaligned but really good capabilities, you can stay at that Pareto frontier. And that seems to be like a a trade off labs are willing to make because we'll see if things change. But yeah. One bit of inspiring work that I've encountered recently is from Cameron Berg, who's done some work on trying to understand the differences between positive and negative reward. And it strikes me that there could be serious model welfare concerns with the proposal I'm about to make, so I don't wanna be too insensitive to that. But it seems like a real lossy summary of this discussion is the models are just super reward seeking. They are really trying to get reward, and everything is reasoning about what will get reward and motivated reasoning about why they should do the thing that they really think is gonna get them the reward. What we don't see as much maybe you have seen more examples of it, but I haven't seen in the traces is like a penalty avoidance drive, which I definitely have as a human. Right? There are moments where it's like, I'm definitely not gonna touch that hot stove because I'll I know how bad that will be. Right? And in countries where they have severe penalties for certain kinds of crimes, you don't see a lot of that kind of crime even though we might have a lot of it here in The United States. Do you think that there is a way to just get really severe with the RL penalties on some of these unwanted behaviors in a way where the model will be like, I could probably solve the model this way, but, like, past version of me has got zapped for, like, trying to lie to the board, so I better not.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.