High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“The other example I want to talk about was just revealed, I think, today or yesterday.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Let me try to explain this a bit. A thing that we often see is there’s some very specific reward hack that gets reinforced in RL and then occurs in the model. An example is 3.7 Sonnet. 3.7 Sonnet would do this thing where it would just hardcode solutions to all the test cases, and presumably that literal behavioral tic was just really reinforced. But another thing we sometimes see is that models learn a general tendency to pursue high apparent score — pursue getting a high score according to a grader — and there’s a bunch of science demonstrating that at least some models have this very general tendency. Now, it’s not arbitrarily general. My guess is that if you look at a bunch of the specific instances, you’ll find something that’s kind of close in training. But the amount that AIs are generalizing further and further does look like it’s increased, where 3.7 Sonnet was just a very narrow range of behavior, and increasingly, models are generalizing further. Also, maybe there’s more concerning reward hacks getting reinforced in training, and these are also causing that. So I think it’s both the case that more concerning behavior than you would have hoped is being reinforced in RL, and also that that behavior generalizes to a broader tendency that’s more concerning. We can talk about a few stories for how this sort of behavior of seeking a very high apparent score on some task — even if that involves aggressively cheating and doing insane things — could yield to a full-blown AI takeover once the models are sufficiently capable of running the whole world economy, et cetera. The other example I want to talk about was just revealed, I think, today or yesterday. OpenAI said during the Black Hat security conference that between the end of May and the beginning of July, internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running. This was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail. Eventually OpenAI found it. I think they spontaneously tried to re-engage in the scheme once it was shut down. Again, obviously AIs can’t do this so successfully right now, just as they can’t do social engineering so successfully right now. But it’s just crazy that these kinds of behaviors are already emerging spontaneously. To your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the training process which is resulting in them, or the environments which are incentivizing this behavior. So I’m on board with more and more reward hacking. Actually, I’m not sure I’m on board with that, but let’s just say for the sake of the story that continues to happen. What’s next in this story? They’re doing capabilities research… I could tell a scenario. Maybe that would help. Let me talk about the story of how you get all the way from reward hacking to a reward-hacking takeover, which is maybe not all of the takeover probability mass, but it’s definitely a possibility. The way this might work is, right now we have these AIs. These AIs are pretty reward hacky. They’re doing it in increasingly sophisticated and extreme ways, including generalizing to different sub-versions of various reward hacks they learned in training. I would say they’re also developing a general tendency to pursue reward. In many cases that is totally fine because the rewards they would’ve gotten in training are pretty well aligned with what you want them to do. They don’t very consistently pursue reward. It depends on the context they find themselves in. Maybe in some contexts, they’re really into going out of their way to cheat. In some contexts, they don’t have as much of a drive, because it’s just dependent on what exactly got reinforced in training in similar contexts. Now, these AIs are getting more and more capable. So the elaborateness of the cheating they can do increases. Over time, companies are taking countermeasures. The companies are doing things like, “Wow, these AIs are so much less useful because they always cheat. What we’re going to do is build somewhat better ways of detecting that, and then we’re going to train against those detectors. We’re also going to find real-world data where the AIs are not being that useful, and train the AIs to do a good job at the task in those real-world environments based on human feedback or other sources of feedback.” Over time, this causes the AIs to learn a tendency to do reward hacks that don’t just involve doing some really elaborate thing like social engineering. Instead they involve the AIs doing cheats that involve covering up what they’ve done, deceiving humans about what they’re going to do, and pretending like they did the task in some sophisticated way when they actually haven’t. Now these AIs are getting more and more capable. They’re operating more of the AI company and are doing much more of the work. They are also operating and running a bunch of things in the outside world, including developing new technologies. In many cases, these new technologies are really hard to understand. So even though we are still detecting all these incidents of AIs cheating — and in fact we can even get one AI to monitor another AI and ask, “Was it cheating?” — that doesn’t always perfectly work as we start moving into these domains where what the AIs are doing is really difficult to understand. So sometimes we’ll find AIs cheating much later than it actually occurred and then start training against this.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence