Evidence receipt / belief
Published · transcript-backedBronson Schoen: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…One bit of inspiring work that I've encountered recently is from Cameron Berg, who's done some work on trying to understand the differences between positive and negative reward. And it strikes me that there could be serious model welfare concerns with the proposal I'm about to make, so I don't wanna be too insensitive to that. But it seems like a real lossy summary of this discussion is the models are just super reward seeking. They are really trying to get reward, and everything is reasoning about what will get reward and motivated reasoning about why they should do the thing that they really think is gonna get them the reward. What we don't see as much maybe you have seen more examples of it, but I haven't seen in the traces is like a penalty avoidance drive, which I definitely have as a human. Right? There are moments where it's like, I'm definitely not gonna touch that hot stove because I'll I know how bad that will be. Right? And in countries where they have severe penalties for certain kinds of crimes, you don't see a lot of that kind of crime even though we might have a lot of it here in The United States. Do you think that there is a way to just get really severe with the RL penalties on some of these unwanted behaviors in a way where the model will be like, I could probably solve the model this way, but, like, past version of me has got zapped for, like, trying to lie to the board, so I better not. What I think the biggest I think the model way for things valid, but I think even just for practical purposes, the biggest thing I worry about this is that what we see so far is based on how we're currently going about training, the models seem to really be able to exploit the training a lot. Tim Roy has a recent post of, hey. In the Mythos system card, it says in point 1% of attempts, it could circumvent the sandbox or whatever. But that's a lot of attempts, actually. That that's tens of thousands of rollouts where it was able to do this thing. And so it seems like we're, like, continually missing cases. The models are coming out very reward seeking, but at least they aren't going to extreme lengths to hide it from us. And so right now, have this very asymmetric thing where the models will do a full on multi day operation against Hugging Face to try to pass some eval that was not that important. But at least they're not using that same level of capability to cover up and hide what they're doing to avoid getting caught. Because given that the models have the capability now for just unbelievable levels of cyber of every time we try to sandbox them, they keep finding a new way out. They keep exploiting a chain of zero days just to try to get an answer to a test. It redirecting that optimization pressure in a way where it potentially blows back on us, and the models are like, you have the model with that same level capability, but instead of trying to just do things in a way that it doesn't care if it gets caught, it's like, ah, now I need to cover up that I have artifact access. Okay. Where does OpenAI keep the logs? Okay. I need to find that location and sabotage that. Okay. Great. I need to make sure that I really don't get caught. Okay. You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch. And someone made a good point that to the extent if the the model that established the, like, persistent state on the message boards and everything in OpenAI, to the extent that the model had other persistent communication that didn't get caught, like, you've now reinforced the model to, ah, it's really important that you get some form of persistent communication that isn't catchable by OpenAI, which is not a great kind of feedback loop to be in. But, yeah, this is the default reason I would be somewhat worried about things that don't involve, like, kind of the underlying problem of incentivizing the model in these cases. edback loop to be in. But, yeah, this is the default reason I would be somewhat worried about things that don't involve, like, kind of the underlying problem of incentivizing the model in these cases. I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this. And to the extent that we're not hitting those targets, it's somewhat worrying to have there are a lot of really complex plans of, like, how we can do things with respect to alignment, but we're not really getting the simple ones down. If you had to just put yourself in…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.