High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

David Rein: recommendation

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“yeah, I think you have to measure the model's or the system's ability to generalize to kind of actually novel situations.”

— David Rein

Source trail

Everything needed to verify it.

Speaker
David Rein
Attribution
Verified speaker
Claim type
recommendation
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…human intelligence is working? We could think of intelligence in many different ways. So is it a simulacrum of the brain? Is it something that behaves the same way? Is it something that has the same capabilities? Is it something that has the same function? And guess if we have quite an abstract description of what intelligence is, the risk is that we have these shortcuts, right? That it might give us the right answer, but actually it's reward hacking or it's doing something silly in the background. So I mean, in a way, I like having an abstract thing, right? Because it's legible, we can evaluate it and so on. But doesn't that leave this kind of hanging risk that it might not actually be doing the thing? And I guess yeah, I think you have to measure the model's or the system's ability to generalize to kind of actually novel situations. There are cases where it seems like models are generalizing well. I think there are cases where they're not. 1 thing some folks do in interpretability is they look at the circuits in models and decompose exactly the algorithms that models are using to answer questions. And sometimes, think, yeah, it seems like they're using shortcuts. Sometimes it seems like they are finding robust patterns. Although, of course, I don't think that work is developed enough to explain most of their behavior currently. But I totally agree that, yeah, you you do have to be pretty concerned with, like, yeah, how how well they're generalizing. You know, operationalizing, what do we care about in the definition of intelligence is that it, like, you know, allows us to predict, like, how will the moles affect the world and what you know, you know, predict what will happen and know how to how to handle them well and things. And so, like, if you just do the black box thing, you you know, maybe that will give you something that doesn't have good generalization because you, like, thought that it was a measure of some type of ability, but it's actually being hacked or a shortcut in in some way. So I think, ideally, what you'd want is, like, generalizing to your benchmark is, like, the same distance as generalizing to the real world from the training data. And, like, that's the sort of thing we thought about, like, when we're doing elicitation on a subset of, the benchmark, we want that, you know, the sort of gap between that subset and the rest of the benchmark to be similar to the gap between that the rest of the benchmark and the real world. And I think we're, like clearly, the training data is more similar to the, you know, to the, like, time horizon, suite than than they both are to, sort of randomly selected economically relevant tasks in the real world. So I think we're not, you know, that's a way in which we expect it to not be predictive. But I think I expect to be more promising to try and make things more predictive by increasing the diversity of the benchmark tasks and making them closer to the real world as opposed to sort of targeting a more mechanistic, like, okay. Intelligence has to be, like, these specific this specific kind of process or or or kind of mechanism.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence