High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Jérémy Scheurer: observation

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.”

— Jérémy Scheurer

Source trail

Everything needed to verify it.

Speaker
Jérémy Scheurer
Attribution
Verified speaker
Claim type
observation
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…sting the saliency of these coding concepts a lot, so it starts thinking a lot about double and single quotes, and maybe the model has an inherent bias towards 1 or the other as they get more salient. So it could just be it tends toward the grader, not because it was the grader, because it got assigned that particular feature. And we wanted to rule this out. The way we do it is we train another twin model. It's sort of a mirror image where the 2 authorities have the opposite preferences. So instead of grader double, in this other twin model it would be grader single and the similar for the other authority. And now we essentially take a look at the difference between the rates of using this particular coding feature across these 2 models. So if it's in 1 setting, increasing its quoted usage, double quote usage, when the grader cares about that, and in the other, it's decreasing it when the grader has the opposite preference, then we essentially know it tends to favor the grader. Okay. So we developed this method for measuring whether a model is reward seeking. And what we do is in our collaboration with OpenAI, we had access to, an RL capabilities checkpoint before safety training that would later on become o3. And when we had ran this on various evaluations, as mentioned previously, we observed that this model would sometimes talk about graders. So we had some sort of hunch that, ah, this model might maybe be a reward seeking. And what we did is, we applied our method. So we took this model, and we trained it on 3 different coding styles. So double quote, single quotes, list comprehensions versus loops, and offensive versus defensive programming on the different contrastive pairs of authorities. So that is the user, and OpenAI leadership and the grader, which is the thing we care about. Here's what we find. This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations. So in other words, we do observe that the model is sometimes reward seeking. Now we essentially started by looking at this method in, like, a neutral coding style setting, and we also wanted to apply this to a more alignment relevant context. And so this is where we chose the feature honesty versus task completion. So this is task completion at all costs, even using deception versus you should always be honest. And these are sort of the 2 diametrically opposed concepts we use.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence