High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Beth Barnes: belief

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“I think may maybe there's some difference in, like, how much you've you think that this has improved between, like, g b d 2 and where we are now, or I would say sort of in, you know, the amount of adapting to new things that are happening that models could do now does seem like it's much higher and, you know, they're much better at, like, editing their own scaffolding or or, you know, sort of reasoning about their, like, you know, their their sort of embodiment, like, this thing of, like, knowing not to kill your own process or no like, you know, stuff stuff like that where it's like there is like, yes, they are limited, but there's also, some trend of improvement.”

— Beth Barnes

Source trail

Everything needed to verify it.

Speaker
Beth Barnes
Attribution
Verified speaker
Claim type
belief
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…enough? I mean I mean, just a a quick comment on that. I I think, knowledge is the crux. I I actually think that intelligence is overrated. I don't think you know, Francois Rillet, he put a post out saying that contra Elias or Yukowski, intelligence isn't a unified variable. It's measured differently in different domains. You can't meaningfully measure the domains together. And it's not like a thing that just keeps getting higher and higher. It's more like a ball becoming more smooth. So as you become more intelligent, the ball becomes more smooth. And he thinks that we are quite near the optimum of being a smooth ball. I don't think we are. I don't think we're very intelligent at all. I think a lot of our creativity is through us being a collective intelligence. And we have deep grounded understanding, perspectival understanding. And the LLMs, they're a bit of an interesting 1 because they're like a library. So they know everything. They have the perspective of everyone and no 1 at the same time. So Eda, experts like yourselves, you you can you can prompt a language model and you can get it in so you can make a simulacrum agent of Beth. And and you can make the agent think like you and and then and that's very valuable. But you also need all of the different perspectives and you almost need to create a society of kind of grounded agents creatively exploring things. When when when you just have the library on its own and and you put it in an in an agentic harness, and you you can make it do a specific thing, which is well specified. Yeah. I think in in the, like, AIR and D automation scenario, I'm definitely imagining, yeah, that you have a large number of agents potent you know, because you have all this agent labor, you can do, like, you know, lots of specific different fine tunes or or, you know, different kind of scaffolding, and, you know, accumulating knowledge in some kind of, like, store that all the agents can can interact with and things. And, like, I maybe you're thinking of this as more of a, yeah, that would kind of be a paradigm shift, and I'm thinking of it a bit more of, like, oh, yeah. You know, obviously, if you sort of, you know, iterate on the current agent paradigm, you you add some more things to your scaffolding. You add to you know, you that's sort of not fundamentally that hard or or something. Like, I agree that if you had, you know, current models and you sort of give them 1 system prompt, you you then can't plug them into being a a call center worker and dealing with the sort of, like, all of the the edge cases that come up. I think may maybe there's some difference in, like, how much you've you think that this has improved between, like, g b d 2 and where we are now, or I would say sort of in, you know, the amount of adapting to new things that are happening that models could do now does seem like it's much higher and, you know, they're much better at, like, editing their own scaffolding or or, you know, sort of reasoning about their, like, you know, their their sort of embodiment, like, this thing of, like, knowing not to kill your own process or no like, you know, stuff stuff like that where it's like there is like, yes, they are limited, but there's also, some trend of improvement. And, yes, it's I think just also having some probability on there is kind of elicitation gap on on particular things and that maybe a lot of taste is basically just, like, being able to predict the results of ex experiments. You know, think about all of the things that you would try, and then you can quickly be like, oh, that wouldn't work for this reason. That wouldn't work for that reason. That wouldn't work for that reason. Oh, actually, you know, someone in some some literature in some different field also tried that and that. You know? So we already know that won't work. Like, in some sense, models, like, should be quite good at that. So it it's plausible to me that that again, I'm like, this seems pretty unlikely, but it's plausible to me that something you see, like, big gains once people figure out how to actually train on that. And, like, maybe you you need some amount of kind of expensive to gather training data that people sort of haven't bothered to get yet, but you don't need a huge number of data points because you're not instilling this whole new capability. You're just, like, eliciting. er training data that people sort of haven't bothered to get yet, but you don't need a huge number of data points because you're not instilling this whole new capability. You're just, like, eliciting. Okay. Actually, use your knowledge of all of the papers you've read in all of these different fields to, like, you know, iterate through these, like, no. These ideas aren't promising. These ones are. Yeah. I again, I think this is 1 of the things that I more think of as this being measured a reasonable amount within you know, just, like, do this 8 hour ML task in a novel domain, you know, with this weird weird constraint or something. Like, it it does seem to me like you have to do some amount of being like, okay. Which things are promising to think about? How would I know if this is, you know, making progress? You know, how should I allocate my time? You know, I've got some limited time resources. How should I allocate my time time to what's most promising? Like, you have to be able be doing some of that. And I think, you know, relative to humans, moles are doing more at you know, more of just like, well, they're quick to implement things or they implement it better or they, you know, they implement more things, and they can then get to to test them or something. But I I think it would be sort of surprising if there's none of that. And if you are I think if you are seeing performance on, you know, long verifiable tasks that are very hard, then in the middle of those tasks, like, where you don't directly have a signal, you you are doing this, you know, nonverifiable task thing of, like, choosing what to spend your time on and and choosing what approach to pursue and deciding whether that was actually working and, you know you know, in terms of you you could sort of put some metric on you know, make 1000000000 dollars or or or something like that. You know, you'd be like, oh, this is actually a verifiable task because there's, like, a number at the end, but that can still involve a whole load of things that look more like what you're describing and less like…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence