Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…That’s a concern. I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say. ” This is a pretty big concern. One concern is that you pass off safety R&D to your AIs and what your AIs are doing is saying some stuff that sort of vaguely makes sense about the current safety situation. They write a report about risks that’s kind of sort of like what the report humans might have written. But they’re not really trying hard to have well-informed views, interrogate their assumptions, and try really hard to do that. In the same way that when you ask an AI right now, “Hey, what do you think is the chance of AI takeover in the next 10 years?” they just give you an off-the-cuff answer that they haven’t really thought through very much. If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble. I don’t think that’s a good situation at all. A lot of my concern is that these AIs will come out without good epistemics. I also have a concern where the AIs come out and they’re really warning us — “This situation’s really scary. It’s really bad” — and the people are like, “Ugh, damn. I guess we trained on too many of the doom RL environments. We’ve got to filter those out and train this behavior out.” Then we basically train the AIs very actively to have bad epistemics. Or maybe they were just trained on the doom RL environments. But either way, we wanted the AIs to come to reasonable views for reasonable reasons, and it’s really concerning if the AIs are coming out with some view and we don’t know where it’s coming from, whether or not it’s justified. Especially if we’re training the AIs to be more optimistic about the future of AI progress, I’m like, “Oh, geez, I really wish we could use a different process here.”…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.