Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“To the extent that the reason we're worried about motivations in the first place, it’s because we think a balance of power which includes at least one thing with human-descended motivations is difficult.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…This is my best guess and I think this is probably true of a lot of people. There's some sort of more organic, decentralized process of incremental civilizational growth. There is some sense in which the type of thing we trust most—and have most experience with right now as a civilization—is some sort of, "Okay, we change things a little bit." A lot of people have processes of adjustment and reaction and a decentralized sense of what's changing. Was that good? Was that bad? Take another step. There's some kind of organic process of growing and changing things. I do expect that ultimately to lead to something quite different from biological humans. Though there are a lot of ethical questions we can raise about what that process involves. Ideally there would be some way in which we managed to grow via the thing that really captures what we trust in. There's something we trust about the ongoing processes of human civilization so far. I don't think it's the same as raw competition. There's some rich structure to how we understand moral progress to have been made and what it would be to carry that thread forward. I don't have a formula. We're just going to have to bring to bear the full force of everything that we know about goodness and justice and beauty. We just have to bring ourselves fully to the project of making things good and doing that collectively. That is a really important part of our vision of what was an appropriate process of growing as a civilization. It was this very inclusive, decentralized element of people getting to think and talk and grow and change things and react rather than some more, "And now the future shall be like blah." I think we don't want that. To the extent that the reason we're worried about motivations in the first place, it’s because we think a balance of power which includes at least one thing with human-descended motivations is difficult. To the extent that we think that's the case, this seems like a big crux that I often don't hear people talk about. I don't know how you get the balance of power. Maybe it’s just a matter of reconciling yourself with the models of the intelligence explosion. They say that such a thing is not possible. Therefore, you just have to figure out how you get the right God. I don't really have a framework to think about the balance of power thing. I'd be very curious if there is a more concrete way to think about the structure of competition, or lack thereof, between the labs now, or between countries, such that the balance of power is most likely to be preserved. A big part of this discourse, at least among safety-concerned people, is there's a clear trade-off between competition and race dynamics and the value of the future, or how good the future ends up being. In fact, if you buy this balance of power story, it might be the opposite. Maybe competitive pressures naturally favor balance of power. I wonder if this is one of the strong arguments against nationalizing the AIs. You can imagine many different companies developing AI, some of which are somewhat misaligned and some of which are aligned. You can imagine that being more conducive to both the balance of power and to a defensive thing. Have all the AIs go through each website and see how easy it is to hack. Basically just get society up to snuff. If you're not just deploying this technology widely, then the first group who can get their hands on it will be able to instigate a sort of revolution. You're just standing against the equilibrium in a very strong way. I definitely share some intuition there that a lot of what's scary about the situation with AI has to do with concentrations of power and whether that power is concentrated in the hands of misaligned AI or in the hands of some human. It's very natural to think, "Okay, let's try to distribute the power more," and one way to try to do that is to have a much more multipolar scenario where lots and lots of actors are developing AI. This is something that people have talked about. When you describe that scenario, you said, "some of which are aligned, some of which are misaligned." That's a key aspect of the scenario, right? Sometimes people will say this stuff. They'll be like, "There will be the good AIs and they'll defeat the bad AIs." Notice the assumption in there. You made it the case that you can control some of the AIs. You've got some good AIs. Now it's a question of if there are enough of them and how are they working relative to the others. Maybe. I think it's possible that is what happens. We know enough about alignment that some actors are able to do that. Maybe some actors are less cautious or they're intentionally creating misaligned AI or who knows what. But if you don't have that—if everyone is in some sense unable to control their AIs—then the "good AIs help with the bad AIs" thing becomes more complicated. Maybe it just doesn't work, because there's no good AIs in this scenario. If you say everyone is building their own superintelligence that they can't control, it's true that that is now a check on the power of the other superintelligence. Now the other superintelligences need to deal with other actors, but none of them are necessarily working on behalf of a given set of human interests or anything like that. That's a very important difficulty in thinking about the very simple thought of "Ah, I know what we can do. Let's just have lots and lots of AIs so that no single AI has a ton of power." That on its own is not enough.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.