High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Ryan Greenblatt: prediction

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
prediction
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Stepping back, I buy the idea that you could have much faster AI R&D than we currently have. I’m not sure if you get GPT-3 to Mythos holding compute and data constant within a year, but suppose it’s half of that. If we even manage to continue the current trajectory of AI progress as a result of AI R&D, it would be fucking insane in five to ten years in ways that I don’t think people appreciate. I don’t think people appreciate what a big deal billions of AIs will be. So I want to understand why you think this might be troubling, Ryan. What could possibly go wrong? What could go wrong? I don’t think we can be so confident about the exact rate of progress here, but it does seem like a lot of rates can be pretty scary. So what could go wrong? Let’s imagine that we’re starting at this point where AI R&D is about to be fully automated or is being fully automated. Things are speeding up, and the way that AI progress is going is kind of crazy. People don’t fully understand what’s going on inside of AI companies. Now, these AIs at the start, they’re not malicious per se. They’re not necessarily very aligned, though. They’re kind of sloppy. They sometimes just do a thing because that’s the sort of thing that would’ve gotten rewarded in training. They aren’t as good at helping you with hard-to-verify tasks due to a mix of poor training incentives — as in, they cheat more or pretend they succeeded when they actually didn’t — and also they’re just less capable at these tasks. But that bites less hard for capabilities, because making AIs more capable has a bunch of verifiable components that the AIs are going really hard at. So then these AIs are getting more and more capable while we understand what’s going on with AI development less and less, and this is happening over a pretty fast period of time. Even just the current rate of progress is, I think, pretty scary. Eventually we get to these AIs that are very superhuman. Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t. Now these AIs are in a position where they’re potentially pretty networked together. They’re operating in neural memory stores that we can no longer decode. They’re thinking thoughts that we don’t fully understand. I think it’s pretty likely that at this point these AIs are scheming against you in a pretty coherent way once they get this superhuman. We can talk about that. Another possibility is that they’re not scheming against you per se, but they are just optimizing for getting a high score on their task. I think that can also lead to AI takeover, which we should talk about. Let’s pause at the first part of the story. So the AIs were not misaligned to begin with, but because the AI R&D is happening really fast, the AIs do end up misaligned? What happened there exactly? I don’t really understand.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence