High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Joe Carlsmith: commitment

22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover

“You might be worried that we will be very gentle and nice and free with the AIs, and then they'll kill us.”

— Joe Carlsmith

Source trail

Everything needed to verify it.

Speaker
Joe Carlsmith
Attribution
Verified speaker
Claim type
commitment
Recorded
22 Aug 2024
Publisher
Dwarkesh Podcast

Transcript context

…The opposite perspective here is that you're doing this sort of vibes-based reasoning of, "Ah, that looks yucky," doing gradient descent on these minds. In the past, a couple of similar cases might have been something like environmentalists not liking nuclear power because the vibes of nuclear don't look green. Obviously that set back the cause of fighting climate change. So the end result of a future you're proud of, a future that's appealing, is set back because your vibes about, "We would be wrong to brainwash a human."You're trying to apply to a disanalogous case where that's not as relevant. I do think there's a concern here, which I really tried to foreground in the series, that is related to what you're saying. You might be worried that we will be very gentle and nice and free with the AIs, and then they'll kill us. They'll take advantage of that and then it will have been a catastrophe. I opened the series basically with an example. I'm really trying to conjure that possibility at the same time as conjuring the grounds of gentleness. These AIs could both be like moral patients—this sort of new species in the sense that should conjure wonder and reverence—and such that they will kill you. I have this example of the documentary Grizzly Man, where there's this environmental activist, Timothy Treadwell. He aspires to approach these grizzly bears. In the summer, he goes into Alaska and he lives with these grizzly bears. He aspires to approach them with this gentleness and reverence. He doesn't carry bear mace. He doesn't use a fence around his camp. He gets eaten alive by one of these bears. I really wanted to foreground that possibility in the series. We need to be talking about these things both at once. Bears can be moral patients. AIs can be moral patients. Nazis are moral patients. Enemy soldiers have souls. We need to learn the art of hawk and dove both. There's this dynamic here that we need to be able to hold both sides of as we go into these trade-offs and these dilemmas. A part of what I'm trying to do in the series is really bring it all to the table at once. If today I were to massively change my mind about what should be done, the big crux that I have is the question of how weird things end up default, how alien they end up. You made a really interesting argument on your blog post that if moral realism is correct, that actually makes an empirical prediction. The aliens, the ASIs, whatever, should converge on the right morality the same way that they converge on the right mathematics. I thought that was a really interesting point. But there's another prediction that moral realism makes. Over time society should become more moral, become better. Of course there is the problem of, "What morals do you have now? It's the ones that society has been converging towards over time." But to the extent that it's happened, one of the predictions of moral realism has been confirmed, so does that mean we should update in favor of moral realism?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence