High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Tim Scarfe: belief

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“I mean, I I I think you you said it was much more common on, rebench than HCAST.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
belief
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…you know, or or, like, you know you know, you're like, I don't know what this thing is doing, but I you know, I'm understanding it as having these goals, and that is useful because I, you know, can make predictions that it'll change the world in certain ways that will result in those goals being achieved. You know, that's kind of how I think about agents. But, yeah. So, reward hacking, in the olden days, you know, people had these demonstrations of reward hacking that were like, the, boat example where it's like, oh, you're you're supposed to, like, go around the track, and they they, like, did some reward shaping by putting coins around the track or something. And then it, like, learned to do some crazy thing where it, like, spins in a circle and catches fire and gets the the coins. And, like, this was, you know, the high scoring thing. And it's, like, in some sense, that's not, that concerning because it's not the the problem is that the the agent is too dumb, and it, like, doesn't have this conception of, like, there was a track and you wanted it to go around the track. It's just, like, doing some pretty blind RL search. So I think that the interesting thing with the more recent reward hacking examples is we're getting to the point where the models smart are enough to understand that that actually is not what you wanted, but they still do it. And you can have a conversation with you know, in chat mode about, like, oh, would you ever do this thing? Or, you know, suppose a user asks you this thing and then you do this. Would that be, you know, aligned behavior? Or suppose some you know, you you know, you can pose it in lots of ways, and, like, clearly, they seem to be able to answer this question of, oh, yeah. No. That was not the desired behavior. But still, they they do it. So I think we've it sort of got to the point where hope 1 hope might be like, oh, the problem was just the systems being dumb. Once they understand what we want, then, you know, you should be able to sort of plug that in somehow to, like, you know, get them to do what we want. But I think it's, like, somewhat interesting that we're seeing it's not trivial to do that even when there is a commercial incentive to do that, which it doesn't mean that we won't, you know, I think it's quite plausible we see the obvious reward hacking being fixed pretty, you know, pretty thoroughly, pretty soon, you know, and sort of a lot of people to tend to say, like, oh, yeah. Yeah. We we just haven't, like, put the best the really good people on it on it yet. It'll it'll it'll get fixed soon. You know? We we once we actually, you know, focus on it, it'll be fine. Yeah. And I'm not sure, but it's it's at least some evidence that it's not trivial to connect the fact that, you know, the model knows this is what not what you want, to it and not actually doing that. I mean, I I I think you you said it was much more common on, rebench than HCAST. And then you can also try and remediate, right? So you can say, please solve this the intended way or you know some people prompt language models that they say kind of like, is we're solving cancer here. This is really, really important that you do it the right way. And some of those remediation prompts actually seem to make it more likely that the model would reward hack. It's a little bit like saying don't press this red button. Right? And then it will press the red button. So how what can we actually do meaningfully to stop this Yeah. I I think, empirically, it seems to happen more on in tasks that are more clearly in the RL distribution rather than the chat distribution on things that have a clear number, and when the agent thinks it's gonna fail otherwise is, you know, sort of the most reward hacky situations. Obvious short term mitigations are to check your RL environments more carefully and read more of your you know, for the for the companies training these models to re like, read what the models are doing more carefully and not reward it for doing things that are obvious hacks. I think the concern there is if you have some detector for reward hacking and you train against it, you may be just overfit to the detector, and you you you're making your reward hacks more subtle, or you train the model to, like, persuade the detector to approve the thing or or or, you know, to it's sort of scary to be in a regime of training against your best ways to, like, know if your problem is is happening because maybe you just get the, like, silent problem. And I think, you know, for the task that for current model capabilities, you know, sometimes it's kind of expensive to have a human check them, but most of the time, it's not beyond any human capabilities. And I think the, you know, the harder version of the problem is when, you know, we're hoping that, like, capabilities will generalize beyond things that we can evaluate, yeah, both from sort of generalization and because we can train on problems even if, like, we wouldn't know how to to solve them or how to, like, look at part of a solution and understand, you know, whether it was sort of doing what we wanted, but we can sort of check the number that comes out at the end. And and and that's a, like, signal that we can use to improve capabilities, but we're gonna be in this regime where you can sort of, like, be superhuman at making numbers go up, but it's unclear…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence