Evidence receipt / belief
Published · transcript-backedJoe Carlsmith: belief
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“I think we as a civilization are going to have a very serious conversation about what sort of servitude is appropriate or inappropriate in the context of AI development.”
Source trail
Everything needed to verify it.
- Speaker
- Joe Carlsmith
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…I want to go back to what our relationship with these AIs should be. Pretty soon we're talking about our relationship to superhuman intelligences, if we think such a thing is possible. There's a question of what process you use to get there and the morality of gradient descenting on their minds, which we can address later. The thing that personally gives me the most unease about alignment is that at least a part of the vision here sounds like you're going to enslave a god. There's just something that feels wrong about that. But then if you don't enslave the god, obviously the god's going to have more control. Are you okay with surrendering most of everything, even if it's like a cooperative relationship you have? I think we as a civilization are going to have a very serious conversation about what sort of servitude is appropriate or inappropriate in the context of AI development. There are a bunch of disanalogies from human slavery that are important. In particular, the AIs might not be moral patients at all, in which case we need to figure that out. There are ways in which we may be able to have motivations. Slavery involves all this suffering and non-consent. There are all these specific dynamics involved in human slavery. Some of those may or may not be present in a given case with AI, and that's important. Overall, we are going to need to stare hard at it. Right now, the default mode of how we treat AIs gives them no moral consideration at all. We're thinking of them as property, as tools, as products, and designing them to be assistants and such. There has been no official communication from any AI developer as to when or under what circumstances that would change. Sothere's a conversation to be had there that we need to have. I want to push back on the notion that there are only two options: enslaved god or loss of control. I think we can do better than that. Let's work on it. Let's try to do better. I think we can do better. It might require being thoughtful. It might require having a mature discourse about this before we start taking irreversible moves. But I'm optimistic that we can at least avoid some of the connotations and a lot of the stuff at stake in that kind of binary. With respect to how we treat the AIs, I have a couple of contradicting intuitions. The difficulty with using intuitions in this case is that obviously it's not clear what reference class an AI we have control over is. Here’s one example, that's very scary about the things we're going to do to these things. If you read about life under Stalin or Mao, there's one version of telling it that is actually very similar to what we mean by alignment. We do these black box experiments to make it think that it can defect. If it does, we know it's misaligned. If you consider Mao's Hundred Flowers Campaign, it’s "let a hundred flowers bloom. I'm going to allow criticism of my regime and so on.” That lasted for a couple of years. Afterwards, for everybody who did that, it was a way to find the so-called "snakes." Who are the rightists who are secretly hiding? We'll purge them. There was this sort of paranoia about defectors, like "Anybody in my entourage, anybody in my regime, they could be a secret capitalist trying to bring down the regime." That's one way of talking about these things, which is very concerning. Is that the correct reference class?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.