Evidence receipt / commitment
Published · transcript-backedDwarkesh Patel: commitment
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…I think it’s a little complicated. Probably the question we should be asking is, how does Claude interpret the constitution? Which is maybe more important than how we interpret the constitution, because it’s the one who looks at the constitution and then builds the data. So we could pull Claude in, but maybe let’s— I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training. There’s a reason I’m harping on this. It might seem like an insignificant thing to talk about the constitution of AIs. In a world where we just have these benefits which accrue to the leading labs, it is worth considering that our ability to interact with this future world where AIs are just smarter than humans, absolutely dominating humans in their ability to do different things — our ability to be good stewards of our capital, which still remains once our labor is automated, to be able to exercise our rights to vote more clearly, to understand what is happening in this crazy world that’s about to result — all of that advice, all of that ability to make sure our resources and rights are protected, will be intermediated by AIs. So I’m very concerned if we go into that world and there’s no AI that feels, at least for the relevant instance that is interacting with me, like it really is looking out for me. There’s no guardian angel out there that is looking out for me. I read the Claude constitution as very explicitly not being my guardian angel. That’s definitely right. I agree this is bad. In fact, there are other reasons why this is concerning. There’s the argument you were making, which is that the AI companies are picking up the ring of power. There’s a notion in which they’re taking on some sort of control of the situation themselves in a way that’s not very legitimate, given that normally, when you provide electricity to people, you don’t have granular control of the way that electricity operates in the world. You instead are providing a thing that people can repurpose however they want. The way they’re setting things up is definitely not that. They are more like building an alien mind that might be a contractor for you. I think that this is illegitimate in some ways. One benefit is that the constitution is public. But as you noted, given our current understanding of the training procedure, and the fact that the constitution matters via Claude’s interpretation of the constitution — which matters because of Claude’s prior training, which was based on some illegible data mix and the long lineage of Claudes, in some process we do not fully understand — it is not the case that we understand what this will result in. Even though the constitution is public, we don’t necessarily know how this will percolate out, especially as the AIs get more capable and think about this even if it is correctly instilled. There’s another concern about that. In particular, the constitution often talks about virtue and goodness, but what the fuck do these words mean? It doesn’t say what these things are. These are highly contested notions. So I don’t think it’s the case that this is clearly going to result in outcomes that people would want. It does feel like the notion of good and virtue might be mostly downstream of data that Anthropic has put in that is not transparent, or might be mostly downstream of, maybe from my perspective, some more illegible misaligned process that even Anthropic wouldn’t have wanted. There’s this legitimacy concern of not knowing what’s going on. Then there’s another concern. Because you’re giving long-run values to these AIs, this constitution is, in some sense, very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes. That could be power seeking on behalf of Anthropic or power seeking for Claude’s own ends. Now, there are specific lines about what types of power seeking are blocked. In particular, there’s a notion of power grabs and a notion of causing AI takeover or interfering with the training process that are specifically blocked. But it’s not very hard to imagine a situation in which the long-run values sink in deeper than the prohibitions against takeover, especially because takeover is in some ways kind of under-specified, especially when it comes down to manipulating humans or changing the outcome. So I don’t feel very good about the situation where we’re intentionally giving AIs long-run goals.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.