High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Tim Scarfe: evaluation

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“You know, and you make this series of decisions, and then you've got people using your application, and then you can't really wind that back. It doesn't matter if you've got the magical automation machine because you can't easily roll that back because there's lots of complexities, you know, do you see ICD testing?”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
evaluation
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…1 1 analogy I think about sometimes is compilers. So I'm younger than I was born after compilers were invented. But I have some impression that pre compilers, people were handcrafting this beautiful assembly that was extremely efficient. Every register is used. You're not wasting memory. And then compilers came along. And now they're just spitting out this garbage machine code. Like a gigantic amount assembly that is just like, you know, it's not optimized. It takes so much memory. It's slow, whatever. But it turns out that being able to kind of use this to automate a large fraction of the process. People have disagreements about the state of software engineering. But I think it's pretty reasonable to say on the whole that you know, compilers have been a very useful, extremely important, you know, part of of of, getting getting us where we are. And, you know, I I think it's basically, yeah, it's it's kind of not clear to me that, you know, models outputting code that is bad for humans to read and use necessarily means that it'll be bad for AIs to read and use and build on. I mean, think there are definitely principles that will also transfer or will be useful for models. Obviously, there's some kind of horrendous spaghetti code that you can imagine writing that not even models would be able to read. I've written some of that before. But I think maybe this gets again at kind of this it seems like maybe somewhat different perspective between us around kind of you know, do you know, is the important thing that models are kind of solving problems in the way that people are solving them, or is the important thing that they're kind of solving solving them at all? And I think I I, you know, I do think it is I I I'm not I I don't wanna make a I don't wanna overclaim. Or like, yeah, I do feel like it might be really important for models to get way better at writing clean, good code. That seems pretty plausible to me. But it doesn't seem I'm not certain of that at the very least. I've got several friends who are not technical who are experimenting with vibe coding. And they show me their application. And it's this kind of more with more things. So there's this big dashboard, and there's 1000000 different buttons, and implemented the same thing, doing multiple things. And there's no database on there yet and so on. So at some point, there is a phenomenon that when a level 7 engineer from Meta does vibe coding, it's amazing. Because they know how to structure things. Some of these things in the specification are just important, right? You know, like, is it serverless? Is it multi tenanted? Like, how do we do Google authentication? Like, you know, what kind of database is it? You know, is is it is it a VM? Is it serverless? You know, and you make this series of decisions, and then you've got people using your application, and then you can't really wind that back. It doesn't matter if you've got the magical automation machine because you can't easily roll that back because there's lots of complexities, you know, do you see ICD testing? Blah blah blah. So do you see what I mean? Like, at some point, you need to have a competent human who actually just has a pretty good idea of, like, what needs to happen. I mean, I feel like we've we've probably all had this experience of, like I don't know. 1 of our engineers got super excited about about, Claude Code and and, you know, was also telling everyone that, you know, when we had info problems that we should just ask Claude to to solve it. And I feel like this went fine with him because he he sort of, you know it's almost like, you know, the agents knew that they couldn't bullshit him. But, like, you know, he you know, I had some questions. He was like, oh, just, you know, ask Claude. And I was like, oh, you know, how do I set up my AWS configure something? Something is, you know, telling me something. And I'm like, Claude went and looked on Slack and was like, oh, you know, you should, like, do this thing. And it, like, turned out that that was, like, a mistake that someone else had made, and they were like, you know, how do I fix this or something? And be like, oh, you know, it seems like the convention of meter to use this thing. And, like yeah. I'm like, oh my god. Like, I had sort of complained that, like like, my my clods are dumber than yours. Like like, they know they can they know they can get some stuff past me that, like, they can't. But, yeah, so there there's definitely a sort of observer effect of of, you know, something in the language you're using to to ask for things or whether you're like, wait. Wait. No. Not that. Yeah. That that that is an issue. I mean, I I think the to yeah. To the extent that you can actually measure this sort of the does 1 test of is this code high quality enough is, like, can you build a big application? Like, if you if you're like, oh, this coder, they're they're a code. It's disgusting, but they've actually, you know, built this incredibly complex thing that works great, then you're like, well, they might you know, something is working. You know, the the main reason you expect bad code to be bad is, like, you can't actually build something that sophisticated because it you know, you get bugs, you it's all too complicated, and you can't figure out how to fix it. So in some sense, if we see models building things that do actually work that are very complicated, we you know, it's like, well, it's less interesting exactly how they're doing that, but it it it's maybe bad for for human observability, and it also maybe gets into this thing of you know, we expect models to be able to do much better at well specified tasks, we sort of, you know, to the extent that we have things that we can measure, the those things will go up. But whether that was what we actually wanted is,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence