High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

David Rein: belief

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“At least me personally, I think we don't have these results aren't like kind of collected and we don't have kind of as systematic results as we do compared to like time horizon partially because, you know, scoring is is is is qualitative now, for for these tasks.”

— David Rein

Source trail

Everything needed to verify it.

Speaker
David Rein
Attribution
Verified speaker
Claim type
belief
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…bunch of reasons why it might not. I think software engineering is a specification acquisition problem. So it's very difficult. We don't know ahead of time what we're building. I'm sure you folks can attest to this, right? So you build some software, and the first version is buggy, and your users use it, and you find lots of edge cases, and then you revise it. And then you have this kind of, thing in your mind, you know, after the tenth revision, and you've kind of created these lovely representations and abstractions and coarse grainings. And you kind of say to yourself, you know what? If I could throw all the code away, could build it 10 times quicker because I know exactly what to do now because I've actually enacted the intelligence. I've actually built the you know, I've found the contours of the domain. It's now an automation problem, basically. And in a sense, contamination thing is a concern for me because when people use Claude code, they're taking your data. So there are people out there that are writing kernel compilers and there are people out there doing all of these different things. And you know Anthropic is just sucking that up and then at some point it becomes an automation problem. So if you're putting a task in there which is essentially a head query so I'm using like information retrieval language here so you know like a head query is it's something that's you know in the mode of the distribution it's used all the time. It's a common task. Claude code will give you the specification because it's already been stolen from other people. Not stolen but you know taken from other people. And then if you give it like something on the long tail then you as the developer have to give it the specification in the prompt. And then again, it's an automation problem. So automation is really easy. So is that what's happening? Like do you think that the increase in in the timelines could just be explained by the acquisition of all of this kind of knowledge from other people doing similar tasks? Yeah. Yeah. I mean, think I think it's a super, a super central question for interpreting where we're at. First thing I'll say, it's hard to know. It's a big question. And so I think we want to have a decent amount of uncertainty. Or we want to kind of take each individual piece of evidence that we've collected as some evidence. We do see models performing better on tasks that have really clear feedback signals, that are extremely well specified, that are in these kinds of domains like in software engineering where you know, if you have written out a spec, you can you can kind of iterate, and and and grind against that. We do also see models performing, much better, on, so called messier tasks, where we haven't already provided this really clean spec. So 1 kind of approach we've taken for creating tasks recently is in particular, yeah, to kind of try and create messier tasks that are less well specified is basically relaxing this kind of automatic scoring constraint. So we don't need to write a really clear, well defined scoring function. And just writing a couple sentences to a model. Hey, build this large piece of software. I'm not going to tell you exactly what I'm looking for. But I'm going to say it needs to be good. And so the model needs to kind of figure out what actually should I build. At least me personally, I think we don't have these results aren't like kind of collected and we don't have kind of as systematic results as we do compared to like time horizon partially because, you know, scoring is is is is qualitative now, for for these tasks. But my my impression is that, you know, models are are worse on these types of tasks than they are, you know, when you give them a clean spec. But they they have been improving at at maybe something like a kind of similar rate. I think there there there are some other sources of evidence we we we have about this, but that's kind of 1 1, major piece of it for me. Messy time. By messy task, mean like ambiguity. And this is absolutely a common thing, right? You know, we do vibe coding and we start off with an ambiguous specification and then reality pushes back and we find the contours of the problem and then we keep telling Claude code, oh actually no, don't do that, do this, do this, do this. And then we we kind of find the shape of the problem and it gets better and better over time. But but but the thing is though, I mean, the the source code for Claude Code leaked yesterday and my friend, he he's a very good software engineer and he was looking for it and and he said, yeah, I don't I don't wanna like, you know, bad talk anthropic. But apparently, it's not very well factored and and, you know, control flow's all over the place. And it's a bit you know, he said if his intern did it, he would have been displeased. But, you know, and I I don't know whether they've even looked at the code. Someone joked actually yesterday there's probably more know, humans are actually looking at the code for Claude code now and maybe they weren't before. But the thing is, there's always like areas of ambiguity and LLMs they do more with more. So intelligence is more with less and LLMs do more with more because the specification the the intelligence comes from the human supervisor. So when you do give them ambiguity, you just get a lot of, like, you know, unfactored code all over the place. So, in in in a sense, like, is is this does that make it harder to evaluate it? Because it might solve it might give you the answer that you're asking for, but it's kind of creating a bit of an unfactored mess at the same time. Yeah. I mean, I think this is a super interesting question.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence