High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Beth Barnes: prediction

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“Like, I think 1 thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't.”

— Beth Barnes

Source trail

Everything needed to verify it.

Speaker
Beth Barnes
Attribution
Verified speaker
Claim type
prediction
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…And she had a really good blog post out recently. And she said that there are 4 big problems, right? So there's like data contamination where the benchmark appears in the training data. Approximate retrieval where the LLMs interpolate from similar training examples without possessing the actual capability to come up with it themselves. Shortcuts, so doing the right things for the wrong reasons. And just more more broadly, not really testing for things like consistency and and robustness and generalization or or the mechanism. It's so much focused just on on the accuracy itself. I mean, how how do you folks think about those kind of problems with benchmark? 1 thing I resonate a lot with there is, you know, the thinking about where are where is most of your error error coming from? And, like, you know, people, like, you know, it it is nice a good good practice to have error bars, you know, based on the, like, standard error in your data or whatever, but that almost always is, like, a tiny fraction of the actual uncertainty. Almost all of it is coming from how does this actually generalize to the real world. So, like, know, a thing we we sort of, like, say to each other a lot at METER is like, but is that the biggest source of uncertainty, or, like, is that the biggest, you know, gap for, like, actually answering the questions we we want to answer? So so thinking about what's the question we're trying to answer? Well, we, you know, we care about sort of things relevant to threat models or or relevant to, like, what the actual impact of AI on the world will be, and therefore, what, you know, properties does our benchmark need to have, or how can we sort of extrapolate across the properties that we can't build in, to be able to make predictions about, you know, the the actual questions that we care about. And I think we think a bit less about the sort of is it doing it the the, like, the real is the model, like, really doing it the right way or something? Like, I think 1 thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't. And I think, like, the sort of history of building those benchmarks has maybe not been amazing. People tend to overfit to those, and you you I think we were trying to have it more be that you capture that thing in that, like if you take a sort of real world relevant, reasonably hard, and long task, then that you know? And you keep that out of the training data, and these tasks are diverse enough at some point if the model is doing that task end to end, it must have had those kind of capabilities as opposed to being able to sort of isolate, you know, a specific sort of theory about it needs to mechanistically be doing this kind of thing. Yeah. I think I think it's interesting because we we have this idea in our minds that humans, we we know how to do things. And when we solve a task that requires reasoning, we kind of follow the specification. We go step by step, and we do things for the right reasons. And when we enact intelligence, we build the specification. We create these course gradings, these abstractions, and they are world aligned. This whole process, that's how we think of human intelligence. And we want the models to kind of behave in that way. Yeah. I mean, think there's an interesting question of whether…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence