Evidence receipt / evaluation
Published · transcript-backedBeth Barnes: evaluation
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“You know, operationalizing, what do we care about in the definition of intelligence is that it, like, you know, allows us to predict, like, how will the moles affect the world and what you know, you know, predict what will happen and know how to how to handle them well and things.”
Source trail
Everything needed to verify it.
- Speaker
- Beth Barnes
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…yeah, I think you have to measure the model's or the system's ability to generalize to kind of actually novel situations. There are cases where it seems like models are generalizing well. I think there are cases where they're not. 1 thing some folks do in interpretability is they look at the circuits in models and decompose exactly the algorithms that models are using to answer questions. And sometimes, think, yeah, it seems like they're using shortcuts. Sometimes it seems like they are finding robust patterns. Although, of course, I don't think that work is developed enough to explain most of their behavior currently. But I totally agree that, yeah, you you do have to be pretty concerned with, like, yeah, how how well they're generalizing. You know, operationalizing, what do we care about in the definition of intelligence is that it, like, you know, allows us to predict, like, how will the moles affect the world and what you know, you know, predict what will happen and know how to how to handle them well and things. And so, like, if you just do the black box thing, you you know, maybe that will give you something that doesn't have good generalization because you, like, thought that it was a measure of some type of ability, but it's actually being hacked or a shortcut in in some way. So I think, ideally, what you'd want is, like, generalizing to your benchmark is, like, the same distance as generalizing to the real world from the training data. And, like, that's the sort of thing we thought about, like, when we're doing elicitation on a subset of, the benchmark, we want that, you know, the sort of gap between that subset and the rest of the benchmark to be similar to the gap between that the rest of the benchmark and the real world. And I think we're, like clearly, the training data is more similar to the, you know, to the, like, time horizon, suite than than they both are to, sort of randomly selected economically relevant tasks in the real world. So I think we're not, you know, that's a way in which we expect it to not be predictive. But I think I expect to be more promising to try and make things more predictive by increasing the diversity of the benchmark tasks and making them closer to the real world as opposed to sort of targeting a more mechanistic, like, okay. Intelligence has to be, like, these specific this specific kind of process or or or kind of mechanism. I'm a huge fan of Francois Cholet, for example. So, you know, he he he created the ARC challenge. And it was just as you say, right? So many, many different tasks. I think 1000 different tasks or so, maybe 800 on the first 1. And they were supposed to be not in the same distribution even though ultimately they were in the same distribution. So distributional leakage was actually the fool of ARC V1 and ARC V2. But the models got really really good at ARC V1 and then Francois released ARC V2 which was different tasks and some of the easier ones were filtered out. And suddenly the LLM performance crashed down to basically 0%. And that to me kind of illustrates that language models, they're good at just seeing many many different examples of things patterns and so on. And and then and then you you you change the task and and they collapse down again. And then ARC v 2 was kind of saturated again 8 months later. So we we do see this pattern. I mean, what do you think about that? With things like ARC AGI,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.