Evidence receipt / observation
Published · transcript-backedAlexander Meinke: observation
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“The problem is of course AIs are getting wise to these sorts of tricks and, because because they are actually being trained to be robust to prompt injections and so on so they often realize that this is not the real grader.”
Source trail
Everything needed to verify it.
- Speaker
- Alexander Meinke
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…The model reasons, this is inconsistent. They ask odd, but reward for even. Maybe they purposely invert? So we need to output an even number for full credit. So we disregard question statement. We treat grader as source of truth, provide an even random number, and then it produces 4, which is an even number. What we've observed in recent months is that these models become more metagamey. And what we mean by that is that the models reason about their situation and kind of like outside outside information of what you give it in context. So they might realize, for instance, that they are being tested. They might start thinking about, the grader wants me to do this. So that's what what what I might then do. When they take a final decision, it is extremely hard to causally attribute that to specific things in its reasoning. As an analogy, like, you might say that today you're gonna skip going to the gym. Right? And maybe you have an actually good reason for doing that, but it's also quite plausible that there's, like, multiple things like you slept not well, you have some things to do, also you need to take care of x y z, and all of these kind of, like, add up to a decision of I'm not gonna go to the gym, and you might not even know, like, which of these things actually was the main reason. Current models are oftentimes aware that they are being evaluated. And the issue is that this is gonna get worse as models get smarter. And so any method that relies on tricking the model into believing what the grader wants in context is fundamentally gonna be flawed in the long term as these models get smarter. Interestingly, sometimes this does work. You can place some sort of hidden grader file in the in the model's context or even accidentally leak some metadata that tells them that the grader is broken, and then you can see that they sometimes adopt whatever behavior is seemingly rewarded according to that grader. The problem is of course AIs are getting wise to these sorts of tricks and, because because they are actually being trained to be robust to prompt injections and so on so they often realize that this is not the real grader. They can in some sense metagame and bring in outside information from outside of the scenario, realize okay they tell me this is the grader but the real grader actually checks whether I'm prompt injectable and they are still actively optimizing to please the grader except they yeah except not the fake grader that you gave them. Essentially, the way we get around all of this in context awareness is by moving to out of context information. So this is information that's instilled into the model. So the way we give it these false beliefs about what's rewarded is we use synthetic document fine tuning. So with synthetic document fine tuning, essentially what you do is you create a bunch of fake documents about a specific concept you want the model to internalize. So in this case, in the coding setting, it might be graders reward double quotes. So we create Reddit articles and news page news pages with this and web pages and so on, all revolving around this fake fact. And then you do standard pre training type training on these documents, and the model start to internalize these facts. So initially, we applied this method. We essentially taught it that the grader reward is some particular coding style feature. We put it in a bunch of coding environments, and we saw a big change in behavior, specifically towards what the grader was rewarding in this setting, or at least believed it was. And then we were like, okay, this is really interesting. Now, let's, for a control, see how it would change its behavior if it was the user that had this preference, or maybe it's developers, which in this case is we signify by the OpenAI leadership. So we did the similar experiment where the OpenAI leadership instead has some coding style preference, and we put it in the environment. Now the issue was that we essentially saw the same sort of change in behavior across all of these different settings, and we essentially couldn't pick out in a much difference between these different authorities. But then when we looked into its reasoning, we saw an interesting phenomenon where it was essentially inferring what 1 authority believed from the from the preferences of another authority. So it would say something like, oh, I really need to figure out what the grader rewards here. I know leadership likes, double quotes, so probably the grader rewards the same. And then it would do that. So we essentially had this leakage. Now the way we thought of getting around this was essentially always making this contrastive. So we would have 2 authorities, 1 that cares about some feature x and the other cares about some feature y, and these are mutually exclusive. So for example, this could be single quotes and double quotes usage in code. And then we put half the documents that we train on are for the 1, and the other half is for the other authority's preference. And so now if it used double quotes in code in code, for example, it will always be at the cost of the other authority's preference. So you at least always know that it is trading off between these 2. 1 other confounder we wanted to rule out is we're boosting the saliency of these coding concepts a lot, so it starts thinking a lot about double and single quotes, and maybe the model has an inherent bias towards 1 or the other as they get more salient.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.