Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, that’s basically right. My perspective is that if AIs are sufficiently good at R&D, including hardware R&D, robots, whatever, then they can radically transform the world, even if they’re not that good at playing politics. Also, we’re in a pretty dangerous situation, because the AIs might be doing huge amounts of really hard-to-understand R&D, building out basically the whole economy of the future, and we may not understand what’s going on in there. Before we move on to the alignment stuff, I think a big source of FUD right now is this realization that this is the way the future is going: extreme economies of scale for the leading labs. The ability to amortize so much intelligence and capabilities across so many different sectors of the economy basically into one model. And not only that, that model will eventually be able to learn from experience. Right now, it’s happening through a process intermediated by humans, where the humans are trying to basically steal your business. They’re like, “Okay, you can do design at Figma, or whatever. We’ll get Claude to do that.” Or, “You can do whatever coding agent. We’ll have Claude internalize that capability.” But eventually, that will be a much more automated process. So there’s this worry that you have models which will basically consolidate all businesses in the world, or at least all current businesses in the world, or at least all current white-collar businesses in the world. Also, at the end of the day, the priority for these companies does not seem to be to release the latest, smartest, most frontier model as soon as they can to as many people as they possibly can. We saw, for example, that Mythos was available internally to Anthropic employees in February, but only released to the public in, I think, June, actually. Also the government got involved, so it ended up being extended almost into July. Between the government and the AI labs themselves, there is this desire to delay the propagation of the latest level of intelligence. Furthermore, there are the concerns about AI takeover, and so we need to solve alignment to make sure there’s no AI takeover. But at the end of the day, there is a real question of: aligned to whom? You look at the way that the constitution of Claude is written. It is just very explicitly not your personal advocate. I’ll pull up some quotes here. “We don’t want Claude to take actions such as searching the web, produce artifacts such as essays, code, or summaries, or make statements that are deceptive, harmful, or highly objectionable. And we don’t want Claude to facilitate humans seeking to do such things. ” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude. ” This is very different from the way lawyers work in America’s current legal regime. Lawyers primarily have the responsibility to help you make your case even if they think you’re guilty. We have decided the way the legal system works best is if everybody has lawyers that are working in their client’s true best interest. There’s not some sense in which the lawyer is really truly motivated by the good of the justice system. em works best is if everybody has lawyers that are working in their client’s true best interest. There’s not some sense in which the lawyer is really truly motivated by the good of the justice system. But I think the way current AIs are shaping up, certainly how Anthropic’s AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end. So there’s this worry that AIs are not, in some deep sense, trying to make sure that I am okay and that my interests are protected in this future, especially given how centralized the development of frontier AI is ending up being. Do you have thoughts on that concern?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.