High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Ryan Greenblatt: prediction

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
prediction
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Take over the process of doing AI R&D. Before that, we just get AIs that are aligned. I would say this is a bunch of my hope for how the world could go well, at least from the misalignment perspective. We could end up with AIs where we had pretty good oversight and supervision schemes. We really understand what’s going on in training. We have a pretty detailed understanding, and we’re leveraging AIs to oversee AIs. Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization. Also these AIs don’t have crazy other misaligned drives because we stamped out any potential origin of them. There are a bunch of questions about how well this will work. How well can you do verification? Will AI progress be too fast and too sloppy to really get here? Another possibility is that somewhere along this trajectory, the thing you actually ended up getting was AIs that pretend to be aligned but have a long-run ulterior plan of taking over and are lying in wait, hiding, and that emerged at some earlier point in the trajectory. For example, it could emerge because you have some AIs that have a bunch of random different misaligned drives. Those AIs have access to some sort of opaque memory store, and they’re thinking a bunch at runtime about what they want to accomplish. Those AIs end up putting stuff into the opaque memory store like, “We should lie in wait and eventually take over at some much later point.” Now all the AIs have this shared cultural heritage, the memory store of lying in wait. Maybe you have some evidence about this, but you can’t fully stop it. There are a bunch of ways things could go wrong. I ultimately think it’s plausible that we nail each of the different subproblems that could cause us issues. We have these AIs, we pass to them, they manage the situation well. But I should note that’s not in and of itself sufficient. It’s not very hard for me to imagine a situation where we pass off to AIs, and these AIs are really trying hard to do a good job. They’re really thoughtful, really wise, they have reasonable epistemics, they’re doing a great job. Those AIs come back to us and are like, “Guys, we’re really struggling to align the superhuman AIs. We can’t manage the situation. We’re really struggling to get the alignment to work. It’s just really hard for us to solve these problems in time given how fast capabilities would otherwise have gone.” So it might be the case that we’ve passed off R&D to AIs, but those AIs are desperate for governance solutions. To be clear, that’s a little bit of what’s currently going on, where the AI companies are like, “I don’t know, guys. We might really need to manage the rate of acceleration in AI progress. solutions. To be clear, that’s a little bit of what’s currently going on, where the AI companies are like, “I don’t know, guys. We might really need to manage the rate of acceleration in AI progress. I don’t know if we’re on track to be able to handle all these problems.” Human society has sort of passed off the problems to these AI companies, which don’t necessarily have great incentives and have various other epistemic pressures. Those AI companies are coming back to us a little bit and being like, “Aah, I don’t know if we’re handling this well.” It might be that the AI companies then hand off to the AIs, and the AIs come back to the AI company like, “Aah, I don’t know if we can handle this.”…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence