High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eliezer Yudkowsky: evaluation

6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

“If you ask me to play a part of somebody who’s quite unlike me, I think there’s some amount of penalty that the character I’m playing gets to his intelligence because I’m secretly back there simulating him.”

— Eliezer Yudkowsky

Source trail

Everything needed to verify it.

Speaker
Eliezer Yudkowsky
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 Apr 2023
Publisher
Dwarkesh Podcast

Transcript context

…Okay, so how about this. I can see that I certainly don’t know for sure what is going on in these systems. In fact, obviously nobody does. But that also goes through you. Could it not just be that reinforcement learning works and all these other things we’re trying somehow work and actually just being an actor produces some sort of benign outcome where there isn’t that level of simulation and conniving? I think it predictably breaks down as you try to make the system smarter, as you try to derive sufficiently useful work from it. And in particular, the sort of work where some other AI doesn’t just kill you off six months later. Yeah, I think the present system is not smart enough to have a deep conniving actress thinking long strings of coherent thoughts about how to predict the next word. But as the mask that it wears, as the people it is pretending to be get smarter and smarter, I think that at some point the thing in there that is predicting how humans plan, predicting how humans talk, predicting how humans think, and needing to be at least as smart as the human it is predicting in order to do that, I suspect at some point there is a new coherence born within the system and something strange starts happening. I think that if you have something that can accurately predict Eliezer Yudkowsky, to use a particular example I know quite well, you’ve got to be able to do the kind of thinking where you are reflecting on yourself and that in order to simulate Eliezer Yudkowsky reflecting on himself, you need to be able to do that kind of thinking. This is not airtight logic but I expect there to be a discount factor. If you ask me to play a part of somebody who’s quite unlike me, I think there’s some amount of penalty that the character I’m playing gets to his intelligence because I’m secretly back there simulating him. That’s even if we’re quite similar and the stranger they are, the more unfamiliar the situation, the less the person I’m playing is as smart as I am and the more they are dumber than I am. So similarly, I think that if you get an AI that’s very, very good at predicting what Eliezer says, I think that there’s a quite alien mind doing that, and it actually has to be to some degree smarter than me in order to play the role of something that thinks differently from how it does very, very accurately. And I reflect on myself, I think about how my thoughts are not good enough by my own standards and how I want to rearrange my own thought processes. I look at the world and see it going the way I did not want it to go, and asking myself how could I change this world? I look around at other humans and I model them, and sometimes I try to persuade them of things. These are all capabilities that the system would then be somewhere in there. And I just don’t trust the blind hope that all of that capability is pointed entirely at pretending to be Eliezer and only exists insofar as it’s the mirror and isomorph of Eliezer. That all the prediction is by being something exactly like me and not thinking about me while not being me. I certainly don’t want to claim that it is guaranteed that there isn’t something super alien and something against our aims happening within the shoggoth. But you made an earlier claim which seemed much stronger than the idea that you don’t want blind hope, which is that we’re going from 0% probability to an order of magnitude greater at 0% probability. There’s a difference between saying that we should be wary and that there’s no hope, right? I could imagine so many things that could be happening in the shoggoth’s brain, especially in our level of confusion and mysticism over what is happening. One example is, let’s say that it kind of just becomes the average of all human psychology and motives.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence