High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Ryan Greenblatt: commitment

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
commitment
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Let me just understand the rest of the threat model, because I think the place where I get off the train is: “Okay, therefore take over the world.” A thing you could imagine is that we just fail to really solve… Let’s just focus on the reward hacking scenario. GPT-8 is making GPT-9. GPT-8 isn’t being super careful. GPT-9 is more “capable” but it is just totally willing to do things like social engineering, hacking, et cetera, but on a qualitatively different scale because it’s a much smarter model. For example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings, if you give it the objective of making a lot of profits this quarter, in a way that causes an Enron-type blowup six months later. Is that the scenario, basically? You have reward hacking, but that reward hacking manifests in companies that are going bankrupt right after the task the CEO is supposed to accomplish is over? All kinds of hacks are through the roof, et cetera. But that doesn’t feel like takeover. That feels more like the equivalent of flash crashes happening all through the economy. Let’s talk about this. I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t. There’s going to be a cat-and-mouse game between AI companies trying to stamp out this behavior and AIs finding increasingly creative reward hacks in training. The equilibrium here is kind of unclear. But one possible outcome is that over time we see increasingly severe and extreme reward hacks — though potentially the rate remains at some intermediate low level — where if the rate of reward hacking gets too high, companies make trade-offs to drive it down. So there’s some equilibrium level where the reward hacking is low enough that it still makes sense to deploy the AI widely into the economy, but high enough that it still causes crazy incidents. Sorry, and this is after GPT-9 has already been deployed?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence