Evidence receipt / belief
Published · transcript-backedJérémy Scheurer: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.”
Source trail
Everything needed to verify it.
- Speaker
- Jérémy Scheurer
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. Yeah. That is something that we will yeah. That we will argue. Yes. It's the capability, also, like, you know, if you just did more pretraining, no RL. Mhmm. You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking. So that's so part of it is yes as it becomes more capable, but more, like, more capable with agentic RL training, for instance. Yeah. I I I think there's 2 separate claims. 1 is that if models are misaligned, then this is worse when they are more capable. Mhmm. But also that your ability to align them might get worse as they get more capable. For this paper, we are more looking at the latter. So the idea that, okay, if models are able to represent the their training process that your ability to keep them aligned to your intent as opposed to whatever their representation of the oversight system is might get worse, but that doesn't automatically imply the other claim.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.