Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
“I don’t know if you saw his recent blog post, but here’s a quote from it: “If you really accept the practical version of the Orthogonality Thesis, then it seems to me that you can’t regard education, knowledge, and enlightenment as instruments for moral betterment.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 6 Apr 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…The broader orthogonality thesis is — you can have almost any kind of self consistent utility function in a self consistent mind. Many people are like, why would AIs want to kill us? Why would smart things not just automatically be nice? And this is a valid question, which I hope to at some point run into some interviewer where they are of the opinion that smart things are automatically nice. So that I can explain on camera why, although I myself held this position very long ago, I realized that I was terribly wrong about it and that all kinds of different things hold together and that if you take a human and make them smarter, that may shift their morality. It might even, depending on how they start out, make them nicer. But that doesn’t mean that you can do this with arbitrary minds and arbitrary mind space because all the different motivations hold together. That’s orthogonality. But if you already believe that, then there might not be much to discuss. No, I guess I wasn’t clear enough about it. Yes, all the different sorts of utility functions are possible. It’s that from the evidence of evolution and from the sort of reasoning about how these systems are being trained, I think that wildly divergent ones don’t seem as likely as you do. But instead of having you respond to that directly, let me ask you some questions I did have about it, which I didn’t get to. One is actually from Scott Aaronson. I don’t know if you saw his recent blog post, but here’s a quote from it: “If you really accept the practical version of the Orthogonality Thesis, then it seems to me that you can’t regard education, knowledge, and enlightenment as instruments for moral betterment. On the whole, though, education hasn’t merely improved humans’ abilities to achieve their goals; it’s also improved their goals.” I’ll let you react to that. Yeah. If you start with humans, if you take humans who were raised the way Scott Aronson was, and you make them smarter, they get nicer, it affects their goals. And there’s a Less Wrong post about this, as there always is, several really, but sorting pebbles into correct heaps, describing a species of aliens who think that a heap of size seven is correct and a heap of size eleven is correct, but not eight or nine or ten, those heaps are incorrect. And they used to think that a heap size of 21 might be correct, but then somebody showed them an array of seven by three pebbles, seven columns, three rows, and then people realized that 21 pebbles was not a correct heap. And this is like a thing they intrinsically care about. These are aliens that have a utility function, as I would phrase it, with some logical uncertainty inside it. But you can see how as they get smarter, they become better and better able to understand which heaps of pebbles are correct. And the real story here is more complicated than this. But that’s the seed of the answer. Scott Aaronson is inside a reference frame for how his utility function shifts as he gets smarter. It’s more complicated than that. Human beings are made out of these are more complicated than the pebble sorters. They’re made out of all these complicated desires. And as they come to know those desires, they change. As they come to see themselves as having different options. It doesn’t just change which option they choose after the manner of something with a utility function, but the different options that they have bring different pieces of themselves in conflict. When you have to kill to stay alive you may come to a different equilibrium with your own feelings about killing than when you are wealthy enough that you no longer have to do that. And this is how humans change as they become smarter, even as they become wealthier, as they have more options, as they know themselves better, as they think for longer about things and consider more arguments, as they understand perhaps other people and give their empathy a chance to grab onto something solider because of their greater understanding of other minds. But that’s all when these things start out inside you. And the problem is that there’s other ways for minds to hold together coherently, where they execute other updates as they know more or don’t even execute updates at all because their utility function is simpler than that. Though I do suspect that is not the most likely outcome of training a large language model. So large language models will change their preferences as they get smarter. Indeed. Not just like what they do to get the same terminal outcomes, but the preferences themselves will up to a point change as they get smarter. It doesn’t keep going. At some point you know yourself especially well and you are able to rewrite yourself and at some point there, unless you specifically choose not to, I think that the system crystallizes.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.