High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Beth Barnes: prediction

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“Like, that that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this, behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, ah, yes.”

— Beth Barnes

Source trail

Everything needed to verify it.

Speaker
Beth Barnes
Attribution
Verified speaker
Claim type
prediction
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…Or you can be reward hacking in a slightly more interesting way where you, like, actually have the goal of making reward go up, you know, and there's, like, planning and stuff going on about that. But these would all be, like, distinct from scheming. Yeah. I I I guess I'm I'm trying to understand the distinction. So you're so you're saying, like, there are examples, you know, like the boat going around, and and that's obviously degenerate behavior. So you you wouldn't interpret that with an agential stance. You would just say that's degeneracy. And when the when the sophistication increases, we might adopt an agential stance and say, it's in service of some bigger goal. But like, I guess the problem I have is, is it always just an interpretation? Could we have a mechanistic or, like, a strong definition of when something is like being an agent? Like, for the specific question that we're discussing, the the test is, like, what does it actually do in some certain circumstance where it has the opportunity to to achieve this this long run goal? So, like, you you can we might not be able to actually observe this, but you can talk about, like, what observations would would make it 1 or the other. So so, you know, like, will this agent, in practice, when it has some opportunity to, you know, make the reward go up, will it do that? Will it will it only do that you know, if the agent is more sort of like the RL algorithm, it'll be like, oh, it will do that once it's, like, explored it by chance and gotten a reward and, you know, that that's been reinforced. If it's, like, an agent that, you know, can reason about the world and planning, it'll be like, okay. It will do that once it, like, you know, learns the learns about the facts about the environment that let it sort of infer that, and then or, you know, if it's we're talking about some, like, long run goal, it's like it would, you know, do it when it actually, you know, has the opportunity to sort of you know, if we're talking about, like, takeover or something, you know, so it's like, yeah, it's not gonna attempt anything while it's under full human control. But once it is, you know, deployed widely enough or has sufficient capabilities to actually succeed in a sort of coup, then it then it would do that. Like, that that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this, behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, ah, yes. It had this other goal, and it's doing what we want and looking like a nice model because it, like, predicts that that will will, you know, lead to it getting more power. And on, Rob's podcast, you you said something that was, you know, quite quite surprising to me. You said that, it you know, AI could autonomously self improve within as little as 2 years and and maybe even shorter timelines were were hard to rule out. Could you, like, walk through the concrete sequence of of steps that could lead to that kind of recursive self improvement? Sure. Yeah. So I think,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence