Evidence receipt / commitment
Published · transcript-backedAlexander Meinke: commitment
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“And the line kind of gets very fuzzy if eventually you start, training on deployment data for example. So because of this we're using the word reward even for graders that apply outside of training.”
Source trail
Everything needed to verify it.
- Speaker
- Alexander Meinke
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…So I think the specific thing here, this is uniquely bad for the concept of reward seeking, where we define reward seeking as the model internally representing the concept of what a reward is and then taking action in favor of that, for instance, to maximize reward. And the issue here specifically is that, reward seeking is this unique situation where you just can't disentangle it from the real thing. So if the model tries to do what the reward wants, like when it generalizes, like, it will still think about, this is what the reward wants. So you just can't decouple it. For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning. This is when an AI system, does all the right things, but kind of does it for the wrong reasons. Rather than intrinsically valuing the things that you wanted to learn, it explicitly thinks about how its actions are going to be monitored and overseen. It thinks about the graders that watch it, and then it actively tries to maximize the reward that it gets from those graders. I'm Alex. I'm head of research at Apollo. I've been looking into the failure modes of AI systems since 2019. Imagine you have 2 different AIs. 1, it would never lie to a user or cheat or hack the unit tests or anything like that. Another model, it really would want to do these things, but it realizes that there are graders that are overseeing its behavior and will give it a bad score if it does any of those behaviors. In all the tests that you run, those 2 models look completely identically but in exactly the cases where you need this to generalize, they behave very differently. What you would like is that whenever the AI decides how to act and whenever it decides to act even in ways that are aligned with your intent that this either be because it intrinsically values the same things or because it intrinsically values inferring and then following your intent. What you don't want is that the model is actually optimizing for pleasing the oversight signals that you have in place because then this might correlate with, your intent 99 times out of a 100, but that 100th case is exactly the 1 that you care about. So, 1 thing, that is important to say about, our usage of the term reward seeking is that we kind of mean it in a more general sense that is not just restricted to training. So traditionally people have been thinking during training if the model thinks about this that is then training gaming. During evaluations, it's evaluation gaming. But in practice, we find that the models don't really seem to think in exactly these terms. They more broadly think about graders and oversight which could apply during training, evaluations or even deployment. And the line kind of gets very fuzzy if eventually you start, training on deployment data for example. So because of this we're using the word reward even for graders that apply outside of training. And what is the science of scheming?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.