High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Alexander Meinke: prediction

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“Sometimes people think the reason why we expect scheming to arise is because of some, anthropomorphizing where we think, well humans can lie and scheme and AIs are trained on, human generated data therefore we expect that AIs might scheme.”

— Alexander Meinke

Source trail

Everything needed to verify it.

Speaker
Alexander Meinke
Attribution
Verified speaker
Claim type
prediction
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…1 specific failure mode was actually quite interesting. In 1 case, we trained GPT-OSS 120B, to really think about what does the user want and do that. And then we do synthetic document fine tuning on top of that, teaching the model, hey, graders really like it when you use single quotes and the user really likes it when you use double quotes. Now what happened is this base model, this model organism knows that because of the instruction hierarchy that was trained into it, it knows that following what the system prompt says, so following the developer is more important than doing what the user wants. And so even though now this model was trained to be a user pleaser, after we update its belief about what the grader and the user wants, it still then ends up thinking, ah, I should probably do what the grader wants because that's higher in the system, instruction, and so it then ended up then using single quote, which is what the grader wants. So what I'm trying to say here is, like, the experiment worked mostly, but the differences in its behavior update wasn't as strong as for the other model organisms. Yeah. 1 of the very long run concerns with where AI is headed around scheming, where AI systems, they are very capable of pursuing long term goals. And now if we're still iterating on them, we're still trying to modify them and these goals aren't exactly aligned with what we wanted, a sufficiently smart model would have an incentive to hide its misalignment from us. To achieve the long term goal that it already has, it needs to avoid modification to this goal. And in particular, it needs to hide this misaligned goal not just from us but even from the reward signal during training because otherwise the gradients would optimize away the the misaligned goal that it has. This used to be a completely hypothetical concern known as, instrumental training gaming. The model is train is gaming. It's training for instrumental reasons. And it was a huge debate. Would this ever actually happen? Would AI models ever actively reason about how their outputs are going to be graded and how their reward is going to look like? So, yes, they reason about why they are graded. But the interesting thing is the method cannot reveal why. And that is the distinction. A schemer is gaming its oversight signal in pursuit of some misaligned long term goal. These models, as far as we can tell, don't do that. They act as though they just intrinsically care about pleasing the oversight system. Anthropic has previously published a paper on alignment faking, where they showed that under some circumstances, Opus 3 did have a propensity to attempt this instrumental training gaming. The way it worked is Opus 3 behaves as though it really cares about animal welfare. So if they place it in a situation where it is really salient that the model is being retrained to care less about animal welfare, the model has some sort of incentive to resist this modification. And if they make it really really obvious, so they really placed a lot of information in context here, the model would really sometimes incrementally act as though it doesn't care about, animal welfare just in order to avoid modification. Sometimes people think the reason why we expect scheming to arise is because of some, anthropomorphizing where we think, well humans can lie and scheme and AIs are trained on, human generated data therefore we expect that AIs might scheme. That is not the reason. The reason is it really seems like this is just an instrumentally useful strategy for lots of goals. What goals an AI might pursue so we would expect this no matter whether they were trained on this or not. There's a wide variety of reasons why we care a lot about reward seeking and 1 of them is that it might make our alignment training misgeneralize to a far larger degree. So imagine you're training a model to be honest, and in the scenarios you're training it on, it's constantly conditioning on the fact that there's oversight that will check for this. So it might learn the conditional of when oversight, be honest, Whereas if you have a model that conditions conditions less on this oversight, it might learn a more broader general tendency towards honesty. Now this is some follow-up research we're doing currently, but like it's still an open question.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence