Evidence receipt / uncertainty
Published · transcript-backedRyan Greenblatt: uncertainty
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.”
Source trail
Everything needed to verify it.
- Speaker
- Ryan Greenblatt
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…What happened with the Hugging Face situation is, in a weird way… I think one of the giveaways to the Hugging Face team… By the way, for context for the audience, Ryan is co-leading the investigation to figure out what happened with the OpenAI Hugging Face incident, so he can’t really comment on this. But I will speculate wildly, because I know this is an opportunity for me to speculate wildly without any rebuttals. It was probably reported that one of the giveaways to the Hugging Face team that this was an AI incident is that the thing was just after this very particular artifact, and not in any other way trying to do something malicious to Hugging Face. So you can imagine a scenario where, let’s say, a deployed instance of GPT-9 is out in the world and given a really hard task: “We want you to design the next great iPhone.” It’s like, “This is so hard. You know what I should do instead? I should just go hack my creators at OpenAI and make sure that in this deployment I’m given a high score.” But then, isn’t the end of the episode that it just hacks into OpenAI servers and gives itself a positive score? Why is it now scheming to get its values into the next generation or something? So one question is, why isn’t it the case that AIs can be really cheaply satisfied by just having some other earlier thing they can hack? You want to succeed at your iPhone task. It turns out you can always succeed by just hacking into OpenAI and messing with them, and then you can just stop there. No need to go further. There’s a few things. One of them is that if this is constantly happening, there might be a bunch of incentive to harden OpenAI. So you’re like, “Fuck it. The AIs keep hacking into OpenAI to mess with their rewards. We’re going to make it so our systems are really, really robust to these AIs hacking in.” Also maybe you start training the AIs to not try to hack into OpenAI in particular. You basically train against each of these specific things. Then one thing you might do is end up selecting for AIs that are more so playing the long game. That’s one concern. Another concern is that your AIs might still be score-seeking, but no longer care about doing that very specific behavior that was very easy, very chill, and now have some broader thing that they ultimately care about. They’re like, “No, no, no, I don’t want to just edit the reward on OpenAI servers. I care about this broader mandate or this broader objective, and I would need to actually make the iPhones.” They actually want to make the iPhones, but they’re willing to take over the whole world to make the better iPhone. That’s another concern you might have. I think it’s kind of unclear exactly how this plays out. But it’s worth noting that if this keeps going on, there’s a bunch of optimization pressure to resolve this. A bunch of the ways it could get resolved are ultimately pretty scary. That’s part of where I’m coming from. Another part of it is that once the AIs are in a position where they can really easily take over the world — we could talk about whether that’s plausible — then I feel like there’s a pretty reasonable case for the AIs. They’re like, “Eh, I don’t know exactly how this is going to go down. I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever. So I’ll both hack OpenAI and, in addition, also take over the world. That will put me in a good position where I have good option value.” If that’s sufficiently easy, the AIs might still do that. Another way to put this is: even if the AIs are pretty cheaply satisfied with some more basic thing, at some point it might just be more reliable for the AIs to take over than it is to just hack into Hugging Face, or even just go to OpenAI and be like, “Look guys, I was able to demonstrate I could steal the answers. Just give me the answers, bro.” Obviously this scenario requires that all this crazy shit is happening. Much smaller incidents keep happening that are still disastrous. Before you take over the world, you cause damage on the scale of billions and tens of billions and hundreds of billions of dollars. Even people die, et cetera. And this does not lead to us solving alignment or shutting down AI development altogether. I just feel like before the takeover happens, society’s just like, “Holy fuck, the AI just killed 1,000 people in order to increase quarterly profits,” or something like that. But maybe this is too much hope that we can at that point be like, “Okay, we have to solve alignment. We have to make sure we know that this thing will not happen again before we keep going.”…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.