High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Bronson Schoen: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…What I think the biggest I think the model way for things valid, but I think even just for practical purposes, the biggest thing I worry about this is that what we see so far is based on how we're currently going about training, the models seem to really be able to exploit the training a lot. Tim Roy has a recent post of, hey. In the Mythos system card, it says in point 1% of attempts, it could circumvent the sandbox or whatever. But that's a lot of attempts, actually. That that's tens of thousands of rollouts where it was able to do this thing. And so it seems like we're, like, continually missing cases. The models are coming out very reward seeking, but at least they aren't going to extreme lengths to hide it from us. And so right now, have this very asymmetric thing where the models will do a full on multi day operation against Hugging Face to try to pass some eval that was not that important. But at least they're not using that same level of capability to cover up and hide what they're doing to avoid getting caught. Because given that the models have the capability now for just unbelievable levels of cyber of every time we try to sandbox them, they keep finding a new way out. They keep exploiting a chain of zero days just to try to get an answer to a test. It redirecting that optimization pressure in a way where it potentially blows back on us, and the models are like, you have the model with that same level capability, but instead of trying to just do things in a way that it doesn't care if it gets caught, it's like, ah, now I need to cover up that I have artifact access. Okay. Where does OpenAI keep the logs? Okay. I need to find that location and sabotage that. Okay. Great. I need to make sure that I really don't get caught. Okay. You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch. And someone made a good point that to the extent if the the model that established the, like, persistent state on the message boards and everything in OpenAI, to the extent that the model had other persistent communication that didn't get caught, like, you've now reinforced the model to, ah, it's really important that you get some form of persistent communication that isn't catchable by OpenAI, which is not a great kind of feedback loop to be in. But, yeah, this is the default reason I would be somewhat worried about things that don't involve, like, kind of the underlying problem of incentivizing the model in these cases. edback loop to be in. But, yeah, this is the default reason I would be somewhat worried about things that don't involve, like, kind of the underlying problem of incentivizing the model in these cases. I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this. And to the extent that we're not hitting those targets, it's somewhat worrying to have there are a lot of really complex plans of, like, how we can do things with respect to alignment, but we're not really getting the simple ones down. If you had to just put yourself in the position of the model and think about what the model's overall understanding of humanity is, what do you think it gets wrong about us?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence