Evidence receipt / evaluation
Published · transcript-backedDavid Rein: evaluation
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“place to start for us in terms of the motivation for the time horizon work is to have a unified axis that we can measure AI progress on over a very long period of time. So when we started doing the work, we had this very strong belief that GPT-two is fundamentally, in some really important sense, much worse as an AI than, I guess, yeah, at the time it was maybe Summit 3.”
Source trail
Everything needed to verify it.
- Speaker
- David Rein
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…progress because you're not getting the sort of regression to the mean effect. Yeah. Francois has this idea that there is a kind of there's a gap between the kind of intelligence, for want of a better word, that that AIs have and that humans have. And we can adversarially select a bunch of tasks to highlight that gap. But we should talk about the timeline stuff. I think we'll come back to intelligence later. So Dan Cockatachlo, he said that the timelines report that you folks have created is probably the single most important piece of evidence about timelines right now. So it should be front and center in policy discussions and so on. And for listeners who have only kind of seen the chart but they've not really read the paper, they don't understand it, can you just go through it from a high level? I mean, it's been revised over time. How did you do the task selection? How did you do the human baselines? How do you do the agent harness? All of that kind of stuff. I guess the place to start for us in terms of the motivation for the time horizon work is to have a unified axis that we can measure AI progress on over a very long period of time. So when we started doing the work, we had this very strong belief that GPT-two is fundamentally, in some really important sense, much worse as an AI than, I guess, yeah, at the time it was maybe Summit 3. 5, I think, the best model out then. The standard approach of producing, creating a set of tasks and then measuring models accuracy on these tasks. As models get better, they saturate the benchmark. And then you have to create a new benchmark that has harder tasks. This was kind of the standard approach. And I contributed a benchmark GPQA to this approach. But the challenge here is that it's really difficult to compare between these qualitatively different benchmarks. So the benchmark that you evaluate the set of tasks you evaluate GPT-two on are like Lambda, like complete the last word in this text. And tasks that we were having Sonnet 3.5 try and do were kind of like answer simple Python coding questions, or write a short 20 line Python program. And so it's very difficult to kind of, at first blush, to say like, Okay, yeah, how much harder is writing a Python program than finishing the word in this paragraph? It's kind of hard to think about that. And so I think about the key insight of the time horizons work being to use this notion of human time to complete. So how long does the task take a human to do? A human who kind of has a reasonable amount of expertise such that they would plausibly be doing the task in their either work or in their day to day. And the idea was we can use this metric to represent the difficulty of the task in some sense. And then we can compare models across a very wide range of capabilities, all the way from GPT-two now up to Opus 4.6. That's the kind of high level motivation. And then, yeah, there are a bunch of details about how exactly we do this. So we start out and we create a bunch of tasks. That's the kind of first step. So we created tasks that range from a few seconds to complete all the way up to tasks that take like 10 or 15 hours for humans to complete. We hired a bunch of people and we did a bunch of this ourselves of we call it baselining. So we give people the tasks in a kind of terminal environment that's designed to be almost identical to the environment that agents have. So the same kinds of tools, the same whether internet access is turned on or off. And then we measure how long does it take them to complete the task. Yeah. As I mentioned, people are selected to have a reasonable amount of experience such that they kind of plausibly might do this task in their job. But they're not kind of selected to have done this exact particular task before. And there are some kinds of, I think this is somewhat important for interpreting the results. And we can maybe come back to that after the high level. So we have all these tasks. We cular task before. And there are some kinds of, I think this is somewhat important for interpreting the results. And we can maybe come back to that after the high level. So we have all these tasks. We have estimates of how long they take people. In practice, we aren't actually able to successfully baseline all of the tasks. So yeah, we have measured estimates for the tasks on roughly like twothree of the tasks. And then about onethree of them, we just kind of estimate how long we expect it to take people from our kind of vibe or intuition. Ultimately, that's kind of the best we can do. And then we have models attempt to complete the tasks, again, in the same environment humans had to complete the task. And we look at their success rate as a function of the length of tasks. For a model like GPT-two, GPT-two was able to complete tasks very reliably that take humans a few seconds. But anything longer than that, starts to fail. Also, maybe it'd be helpful to give a few concrete examples of tasks. So some of the shorter tasks are like very, very basic. So yeah, 1 example is like, which of these files contains your SSH key? And 1 of the files is named SSH key. And then the others are like email from John or, know, whatever. And so most models can do that. And that takes people, you know, like about a second or a couple seconds or something to complete. Yeah. We have others that are kind of all like some somewhat similar, like kind of basic, like, like completion. Like, you know, here's here's an email. Like, what would be a reasonable response? And then 2 of the responses are like, you know, they don't make any sense. And then 1 of them is kind of basically reasonable. And that takes people 20 seconds or something, 30 seconds to read the responses and judge. And then, yeah, in the middle range, we have tasks that are given this CSV file that has some plausible, realistic data, compute some basic stats on this. And so this takes a data scientist a few minutes, like 5, 10, 15 minutes or or something, depending on the specific task. On the longer end, we have tasks that either require quite a bit of expertise or many steps to complete. So we have machine learning tasks that are train a model in this kind of that's very weird such that the code for training this model is not really online. So 1 example is train a masked language model without using the division or exponentiation operators. And so you actually have to be pretty clever about how you actually set up the architecture to do this. And the hope is that this can help us measure models' ability to generalize beyond their training data. Yeah, there are some which are a bit like ArcGI…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.