Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“I don't know if moral realism is the right word, but you mentioned the thing. There's something that makes hearts converge to the thing we are or the thing we would be upon reflection.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…One thing I want to flag is that not all forms of moral realism make this prediction. I'm happy to talk about the different forms I have in mind. There are also forms of things that look like moral anti-realism—at least in their metaphysics according to me—but which just posit that there's this convergence. It's not in virtue of interacting with some kind of mind-independent moral truth, but just for some other reason. That looks a lot like moral realism at that point. It's universal, everyone ends up there. It's tempting to ask why and whatever answer is a little bit like, "Is that the Dao? Is that the nature of the Dao?" even if there's not an extra metaphysical realm in which the moral lives. Moral convergence is a different factor from the existence or non-existence of a morality that's not reducible to natural facts, which is the type of moral realism I usually consider. Now, does the improvement of society update us towards moral realism? Maybe it’s a very weak update or something. I’m kind of like, “Which view predicts this more strongly?” It feels to me like moral anti-realism is very comfortable with the observation that people with certain values have those values. There's obviously this first thing. If you're the culmination of some process of moral change, then it's very easy to look back at that process and say "Ah, moral progress. The arc of history bends towards me.” If there were a bunch of dice rolls along the way, you might think, "Oh wait, that's not rational. That's not the march of reason." There's still empirical work you can do to tell whether that's what's going on. On moral anti-realism, consider Aristotle and us. Has there been moral progress by Aristotle's lights and our lights too? You could think, "Ah, doesn't that sound a bit like moral realism? These hearts are singing in harmony. That's the moral realist thing, right? The anti-realist thing is that hearts all go in different directions, but you and Aristotle apparently are both excited about the march of history.” There's an open question about whether that's true. What are Aristotle's reflective values? Suppose it is true. That's fairly explicable in moral anti-realist terms. You can roughly say that you and Aristotle are sufficiently similar. You endorse sufficiently similar reflective processes. Those processes are in fact instantiated in the march of history. So history has been good for both of you. There are worlds where that isn't the case. So there's a sense in which maybe that prediction is more likely for realism than anti-realism, but it doesn't move me very much. I don't know if moral realism is the right word, but you mentioned the thing. There's something that makes hearts converge to the thing we are or the thing we would be upon reflection. Even if it's not something that's instantiated in a realm beyond the universe, it's a force that exists that acts in a way we're happy with. To the extent that it doesn't exist and you let go of the reins and you get the paper clippers, it feels like we were doomed a long time ago? We were just different utility functions banging against each other. Some of them have parochial preferences, but it's just combat and some guy won. In the other world it’s “No, these are where the hearts are supposed to go or it's only by catastrophe that they don't end up there.” That feels like the world where it really matters. The initial question I asked was, “What would make us think that alignment was a big mistake?” In the world where hearts just naturally end up like the thing we want, maybe it takes an extremely strong force to push them away from that. That extremely strong force is you solve technical alignment, the blinders on the horse's eyes. In the worlds that really matter, we're like, "Ah, this is where the hearts want to go." In that world, maybe alignment is what messes us up. So the question is, do the worlds that matter have this kind of convergent moral force, whether metaphysically inflationary or not, or are those the only ones that matter?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.