Evidence receipt / prediction
Published · transcript-backedBeth Barnes: prediction
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“I give a slightly different number, but, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not un unlikely enough to rule out. And I think that basically looks like maybe we we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit, it on on these less less health and climateable tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing.”
Source trail
Everything needed to verify it.
- Speaker
- Beth Barnes
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…it you know, AI could autonomously self improve within as little as 2 years and and maybe even shorter timelines were were hard to rule out. Could you, like, walk through the concrete sequence of of steps that could lead to that kind of recursive self improvement? Sure. Yeah. So I think, yeah, maybe I'd put, like, a you know, I'm, like, whole whole number percent this year, but but low low whole number percent or something. You you know, ask me on different days. I give a slightly different number, but, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not un unlikely enough to rule out. And I think that basically looks like maybe we we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit, it on on these less less health and climateable tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing. Then this is leading to yeah. Like, you automate and accelerate a bunch of AIR and D. So I think I think there's you know, there are a lot of low hanging fruit even of things that we already know that you could do this and it would improve model performance. And it's just you know, it doesn't require new breakthroughs, and it's just kind of labor intensive to do. So just making much, much better post training environments and really, you know, crafting them to to teach all the new abilities that you want. And and I think you could probably improve compute efficiency a bunch with, you know, again, just, like, applying a bunch more labor to, like, making all your kernels more efficient and also, you know, doing the right kind of of rooting between different models or or or other things like that. There's there's, like, lots of ways in which, how we're using compute is not optimized, so you could potentially get a bunch of, you know, sort of, like, the equivalent of much more compute scaling out of that. And then, you know, high hypothesis, like also, like, scaffolding and training the models to use particular scaffolding and sort of, like, use memory and retrieval in the right way. Like, it seems kind of obvious that, like, you know, if you if you sort of really had all the right training data and you have a, you know, you have a transformer and and it can kind of fill its context with with different things and and and take stuff in and out, it can do a pretty, you know, good job of something. Looks like the sort of continual learning or or or building up understanding if you've got massive massive context window and a know, you've actually got quite a lot of bits in there to sort of be adding things about what you've been learning and and, you know, if you if you really sort of had optimized the training for all of that, like, maybe you can get that to work pretty well. And then your you know, maybe some other piece would be like, yeah. Moles are kind of superhuman at predicting the results of experiments because they've read so many papers and, to work pretty well. And then your you know, maybe some other piece would be like, yeah. Moles are kind of superhuman at predicting the results of experiments because they've read so many papers and, predicting experiments and and sort of synthesizing things from different fields. And, again, maybe this is something like, you know, it's possible that we're not seeing that good performance here just because we haven't quite elicited the models to do it, and it it's, like, not a thing that they've seen humans do, but they actually sort of have the capability in there. Yeah. So maybe you can make much faster progress if you can you can do a bunch of iteration. You don't actually have to run the experiments. Then, you know, models are much better at predicting what will and won't work. And then you can when you do run experiments, you can sort of, run a bunch more of them because you can optimize the code with your, like, you know, very fast coding models. Know? And then you sort of like, as you do a few more rounds of this, you get to a point where train on a bunch more things that are good, high quality task proxies for what you want, and you get enough generalization to the things that you can't, directly train against.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.