Evidence receipt / observation
Published · transcript-backedBeth Barnes: observation
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“there is a sort of adversarial selection going on where people are trying to make some benchmark, like, cheaply subject to the constraint that current models do badly on it, which means you, you know, you can't use a lot of expensive human labor, it has to be something that's either automatically checkable or that you can, like, create with kind of cheap human labor.”
Source trail
Everything needed to verify it.
- Speaker
- Beth Barnes
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…I'm a huge fan of Francois Cholet, for example. So, you know, he he he created the ARC challenge. And it was just as you say, right? So many, many different tasks. I think 1000 different tasks or so, maybe 800 on the first 1. And they were supposed to be not in the same distribution even though ultimately they were in the same distribution. So distributional leakage was actually the fool of ARC V1 and ARC V2. But the models got really really good at ARC V1 and then Francois released ARC V2 which was different tasks and some of the easier ones were filtered out. And suddenly the LLM performance crashed down to basically 0%. And that to me kind of illustrates that language models, they're good at just seeing many many different examples of things patterns and so on. And and then and then you you you change the task and and they collapse down again. And then ARC v 2 was kind of saturated again 8 months later. So we we do see this pattern. I mean, what do you think about that? With things like ARC AGI, there is a sort of adversarial selection going on where people are trying to make some benchmark, like, cheaply subject to the constraint that current models do badly on it, which means you, you know, you can't use a lot of expensive human labor, it has to be something that's either automatically checkable or that you can, like, create with kind of cheap human labor. And then but once you've selected on, like, those things and on models being bad at it, this is now you know, there's, like, regression to the mean type thing where it is much more likely that future progress then gives you a, like, big surge upwards on on that, both because, like, you know, now labs will create a bunch of bunch of synthetic data targeting your benchmark, but also just because you selected this weird example where it's, like, easy for humans or it's automatically generateable or checkable, but somehow, like, models aren't good at it yet or, you know, labs haven't started training on it yet. So I think that's part of what we were trying to do with Time Horizon was not do that, like, not adversarially select against what models can currently do because we think that will not give you a nice trend. Whereas if you can sort of define a distribution of tasks in some more first principles way, you would be more likely to get a steady progress because you're not getting the sort of regression to the mean effect. Yeah. Francois has this idea that there is a kind of there's a gap between the kind of intelligence, for want of a better word, that that AIs have and that humans have. And we can adversarially select a bunch of tasks to highlight that gap. But we should talk about the timeline stuff. I think we'll come back to intelligence later. So Dan Cockatachlo, he said that the timelines report that you folks have created is probably the single most important piece of evidence about timelines right now. So it should be front and center in policy discussions and so on. And for listeners who have only kind of seen the chart but they've not really read the paper, they don't understand it, can you just go through it from a high level? I mean, it's been revised over time. How did you do the task selection? How did you do the human baselines? How do you do the agent harness? All of that kind of stuff. I guess the…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.