Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“Because the whole reason we created agile software development as a methodology is because it's inconceivable, It's outside our cognitive horizon.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Like, yeah, hill hill climbable, yeah, easily checkable tasks that you can do from a terminal or, like, you know, comfortably in a language interface or, like, a text, input output interface. I I think there's a question of you know, do we care about what what is the what are the statements that we're, like, 99% confident in? We maybe are also interested in the statements that are we're 1% confident in if we're like, oh, there's, like, like, 1% chance that we have a crazy intelligence explosion, you know, in '20 at the the 2026 and that, you know, sort of the fate of civilization depends on, like, how that goes or something. You know, that is that is interesting to know even if it's a even if you're 99% confident that it won't happen. Like, we we you know, we care about things that are 1% you know, if you have some some diagnosis, it's like, it's a 1 percent chance that you have this, you know, terminal illness or something. You're you're still like, oh, shit. So I, yeah, I think we're interested in, like, the whole distribution of, like, what things can we rule in and what things can we rule out and what things are we like, oh, actually, you know, that there's a kind of reasonable story for this. It seems like probably, you know, pretty unlikely, unlikely, but, but, like, like, maybe this is now in the realm of, like, we should consider it. You probably saw the the the Carlini paper for is is at Anthropic now. And they got a swarm of agents to create a compiler. And, in a sense, I mean Jeremy Howard, when I spoke to him, he said it's basically a style transfer problem because the specification is online and the tests are online and the code is online and it could just iteratively just do the thing until it worked and then it could run Doom and all this kind of stuff. That is an example of an extremely complicated piece of software because I often joke to people that, you know, the best mark of AGI is when it could build something like the Linux operating system. And I guess just like in line of what we were saying before, we have this specification problem, right? So it gets to the point where no human could understand or create the specification for the Linux operating system. So what happens is that over time, we've just kind of incrementally built this thing because our brains are limited. We take 1 step, reality pushes back, we take 1 step and we just keep going and we build the specification. But what would it mean to, as a human, specify a task that could take 4 months? Because the whole reason we created agile software development as a methodology is because it's inconceivable, It's outside our cognitive horizon. So isn't that a bit of a chicken and egg problem that in like in my mind, I don't think we could specify a task of that complexity. Therefore, the AIs wouldn't be able to do it. 1 analogy I think about is, you know, the role of a CEO at companies. Actually, yeah, maybe Beth is better put to answer this. But CEOs do come up with a vision for where they want the company to be. And then they communicate that concisely, their executives that report to them. Then if they're a good CEO and if the company is effective, then the company is able to kind of take this very concise it's not actually that much information. It's not the full spec at all. It's not even close. Right? And turn that into something that is aligned with what they're looking for. And so there's at least this is, to some extent, we do have examples of people being able to specify some task and then be able to judge whether this very large task that may take hundreds or thousands of person years to actually complete because it requires many people working over a long time. And they can judge whether they've succeeded or failed. So that's maybe 1 motivating intuition where language you know, has built in or or like, you know, we do we do have like kind of, you know, you know, and there there is kind of enough meaning or or expressivity or something to be able to kind of have some some kind of reasonable understanding. Obviously, are like, tons of edge cases. And often CEOs aren't able to get their companies to do what they want. But yeah, I guess that's 1 1 thing I think about. And and so that's why I think it's at least plausible that AIs could, could could could do these kinds of long tasks. I would say that saying something like we can't specify tasks that take more than 4 months seems sort of obviously too strong. Like, there are even, you know, numerical…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.