Evidence receipt / prediction
Published · transcript-backedJérémy Scheurer: prediction
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.”
Source trail
Everything needed to verify it.
- Speaker
- Jérémy Scheurer
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…I guess the interesting thing for me is that when we think of reinforcement learning algorithms like AlphaGo Zero, it makes sense that they are reward seeking because there is this structured inference process. This is a language model. It's doing greedy sampling of tokens. So you're saying that when the models are trained with reinforcement learning, it imprints this reward seeking behavior. Or, I mean, at least that's that's what we think is happening given also our evidence that, like, we've increased RL training, the model becomes or tends to to be more reward seeking. I think the important thing to to kind of, like, reason about this is you mentioned for instance AlphaGo Zero or like a chess engine and I think there you could train this model with as many flops as you want and I don't think the model is ever gonna become reward seeking in any way because it's just like yeah. It's just the the MDP you're working on is basically just like a a a chess engine for instance. Mhmm. But the difference here is that, like, a language model comes pre trained with a lot of these, human concepts, so it has the ontology of the world, and so it knows or it learned what a reward is, and so when you put it into RL, it seems probable that, like, it like, we know, for instance, also that models sometimes know when they're being tested or not, and so it just seems way more, likely that a model could develop reward seeking because it has these kind of concepts internalized. We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that. Now it's kind of enough that at some point, just due to the stochastic process of sampling from these models, the model might output tokens which are like, ah, this is some sort of training run. Maybe I'm being rewarded or something like that. Like, this can just happen, you're right, because we're sampling at high temperatures. The the the thing that could happen here is, like, because this whole trajectory gets rewarded and the model eventually completed a task, that these kind of reasoning traces could get upweighted. It's kind of like a plausible story for how this could happen, I think. When you're trying to get a high reward, it's really useful to think about how are they gonna score me, how are they gonna be evaluating me, how are they checking my answers, and then modifying your behavior to optimize for those. And things that are really useful for getting high reward will get reinforced during the course of training. So it only has to stumble into this sort of reasoning, and that can again get laterally reinforced.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.