Evidence receipt / evaluation
Published · transcript-backedJoe Carlsmith: evaluation
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“The most straightforward thing to focus on—I don't blame people for just talking about this one—is just the first one. It's quite robust according to our own ethics, when we think about in which context is it appropriate to try to exert various types of control, or to have more of what I call in the series "yang," which is this active controlling force, as opposed to "yin," which is this more receptive and open, letting go.”
Source trail
Everything needed to verify it.
- Speaker
- Joe Carlsmith
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…But then one could accuse the alignment community of motte and bailey. The motte is: We just want to make sure that GPT-8 doesn't kill everybody. After that, we're all cool. Then the real thing is: “We are fundamentally pessimistic about historical processes, in a way that doesn't even necessarily implicate AI alone. It’s just the nature of the universe. We want to do something to make sure the nature of the universe doesn't take a hold on humans and where things are headed. If you look at the Soviet Union, the collectivization of farming and the disempowerment of the kulaks was not as a practical matter necessary. In fact it was extremely counterproductive and it almost brought down the regime. Obviously it killed millions of people, caused a huge famine. But it was sort of ideologically necessary. You have an ember of something here and we have to make sure that an enclave of the other thing doesn't put it out. If you have raw competition between the kulak type capitalism and what we're trying to build here, the gray goo of the kulaks will just take over. We have this ember here. We're going to do worldwide revolution from it. I know that obviously that's not exactly the kind of thing alignment has in mind, but we have an ember here and we've got to make sure that this other thing that's happening on the side doesn't FOOM. Obviously that's not how they would phrase it, but so that it doesn’t get a hold on what we're building here. That's maybe the worry that people who are opposed to alignment have. It’s the second kind of thing, the kind of thing that Stalin was worried about. Obviously, we wouldn't endorse the specific things he did. When people talk about alignment, they have in mind a number of different types of goals. One type of goal is quite minimal. It's something like, “The AI's don't kill everyone or violently disempower people.” There's a second thing people sometimes want out of alignment, which is much broader. It’s something like, “We would like it to be the case that our AI's are such that when we incorporate them into our society, things are good, that wee just have a good future.” I do agree that the discourse about AI alignment mixes together these two goals that I mentioned. I actually mentioned three goals. The most straightforward thing to focus on—I don't blame people for just talking about this one—is just the first one. It's quite robust according to our own ethics, when we think about in which context is it appropriate to try to exert various types of control, or to have more of what I call in the series "yang," which is this active controlling force, as opposed to "yin," which is this more receptive and open, letting go. A kind of paradigm context in which we think that is appropriate is if something is an active aggressor against the boundaries and cooperative structures that we've created as a civilization. I talked about the Nazis. In the piece, I talked about how when something is invading, we often think it's appropriate to fight back. We often think it's appropriate to set up structures to prevent and ensure that these basic norms of peace and harmony are adhered to. I do think some of the moral heft of some parts of the alignment discourse comes from drawing specifically on that aspect of our morality. We think the AIs are presented as aggressors that are coming to kill you. If that's true, then it's quite appropriate. That’s classic human stuff. Almost everyone recognizes that self-defense, or ensuring basic norms are adhered to, is a justified use of certain kinds of power that would often be unjustified in other contexts. Self-defense is a clear example there. I do think it's important though to separate that concern from this other concern about where the future eventually goes. How much do we want to be trying to steer that actively? I wrote the series partly in response to the thing you're talking about. It is true that aspects of this discourse involve the possibility of trying to steer and grip. You have a sense that the universe is about to go off in some direction and you need people to notice that muscle. We have a very rich ethical human ethical tradition of thinking about, when it is appropriate to try to exert what sorts of control over which things. Part of what I want to do is that I want us to bring the full force and richness of that tradition to this discussion. It's easy if you're purely in this abstract mode of utility functions and human utility functions. There's this competitor thing with a utility function. richness of that tradition to this discussion. It's easy if you're purely in this abstract mode of utility functions and human utility functions. There's this competitor thing with a utility function. Somehow you lose touch with the complexity of how we've been dealing with differences in values and competitions for power. This is classic stuff. AI sort of amplifies a lot of the dynamics, but I don't think it's fundamentally new. Part of what I'm trying to say is let's draw on the full wisdom we have here, while obviously adjusting for ways in which things are different.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.