High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“Let me just understand the rest of the threat model, because I think the place where I get off the train is: “Okay, therefore take over the world.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…This is a pretty big concern. One concern is that you pass off safety R&D to your AIs and what your AIs are doing is saying some stuff that sort of vaguely makes sense about the current safety situation. They write a report about risks that’s kind of sort of like what the report humans might have written. But they’re not really trying hard to have well-informed views, interrogate their assumptions, and try really hard to do that. In the same way that when you ask an AI right now, “Hey, what do you think is the chance of AI takeover in the next 10 years?” they just give you an off-the-cuff answer that they haven’t really thought through very much. If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble. I don’t think that’s a good situation at all. A lot of my concern is that these AIs will come out without good epistemics. I also have a concern where the AIs come out and they’re really warning us — “This situation’s really scary. It’s really bad” — and the people are like, “Ugh, damn. I guess we trained on too many of the doom RL environments. We’ve got to filter those out and train this behavior out.” Then we basically train the AIs very actively to have bad epistemics. Or maybe they were just trained on the doom RL environments. But either way, we wanted the AIs to come to reasonable views for reasonable reasons, and it’s really concerning if the AIs are coming out with some view and we don’t know where it’s coming from, whether or not it’s justified. Especially if we’re training the AIs to be more optimistic about the future of AI progress, I’m like, “Oh, geez, I really wish we could use a different process here.” Let me just understand the rest of the threat model, because I think the place where I get off the train is: “Okay, therefore take over the world. ” A thing you could imagine is that we just fail to really solve… Let’s just focus on the reward hacking scenario. GPT-8 is making GPT-9. GPT-8 isn’t being super careful. GPT-9 is more “capable” but it is just totally willing to do things like social engineering, hacking, et cetera, but on a qualitatively different scale because it’s a much smarter model. For example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings, if you give it the objective of making a lot of profits this quarter, in a way that causes an Enron-type blowup six months later. Is that the scenario, basically? You have reward hacking, but that reward hacking manifests in companies that are going bankrupt right after the task the CEO is supposed to accomplish is over? All kinds of hacks are through the roof, et cetera. But that doesn’t feel like takeover. That feels more like the equivalent of flash crashes happening all through the economy. Let’s talk about this. I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t. There’s going to be a cat-and-mouse game between AI companies trying to stamp out this behavior and AIs finding increasingly creative reward hacks in training. The equilibrium here is kind of unclear. But one possible outcome is that over time we see increasingly severe and extreme reward hacks — though potentially the rate remains at some intermediate low level — where if the rate of reward hacking gets too high, companies make trade-offs to drive it down. So there’s some equilibrium level where the reward hacking is low enough that it still makes sense to deploy the AI widely into the economy, but high enough that it still causes crazy incidents.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence