High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Jérémy Scheurer: belief

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.”

— Jérémy Scheurer

Source trail

Everything needed to verify it.

Speaker
Jérémy Scheurer
Attribution
Verified speaker
Claim type
belief
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. Yeah. That is something that we will yeah. That we will argue. Yes. It's the capability, also, like, you know, if you just did more pretraining, no RL. Mhmm. You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking. So that's so part of it is yes as it becomes more capable, but more, like, more capable with agentic RL training, for instance. Yeah. I I I think there's 2 separate claims. 1 is that if models are misaligned, then this is worse when they are more capable. Mhmm. But also that your ability to align them might get worse as they get more capable. For this paper, we are more looking at the latter. So the idea that, okay, if models are able to represent the their training process that your ability to keep them aligned to your intent as opposed to whatever their representation of the oversight system is might get worse, but that doesn't automatically imply the other claim.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence