High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eliezer Yudkowsky: evaluation

6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

“I’m too stupid to solve alignment and I’m too stupid to execute a handshake with a superintelligence that I told somebody else how to align in a cleverly, deceptive way where that superintelligence ended up in the kind of basin of logical decision theory, handshakes or any number of other methods that I myself am too stupid to a vision because I’m too stupid to solve alignment. The point is — I think about this stuff.”

— Eliezer Yudkowsky

Source trail

Everything needed to verify it.

Speaker
Eliezer Yudkowsky
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 Apr 2023
Publisher
Dwarkesh Podcast

Transcript context

…I guess if it is as easy to do that, why haven’t you been able to do this yourself in some way that enables you to take control of the world? Because I can’t solve alignment. First of all, I wouldn’t. Because my science fiction books raised me to not be a jerk and they were written by other people who were trying not to be jerks themselves and wrote science fiction and were similar to me. It was not a magic process. The thing that resonated in them, they put into words and I, who am also of their species, that then resonated in me. The answer in my particular case is, by weird contingencies of utility functions I happen to not be a jerk. Leaving that aside, I’m just too stupid. I’m too stupid to solve alignment and I’m too stupid to execute a handshake with a superintelligence that I told somebody else how to align in a cleverly, deceptive way where that superintelligence ended up in the kind of basin of logical decision theory, handshakes or any number of other methods that I myself am too stupid to a vision because I’m too stupid to solve alignment. The point is — I think about this stuff. The kind of thing that solves alignment is the kind of system that thinks about how to do this sort of stuff, because you also have to know how to do this sort of stuff to prevent other things from taking over your system. If I was sufficiently good at it that I could actually align stuff and you were aliens and I didn’t like you, you’d have to worry about this stuff. I don’t know how to evaluate that on its own terms because I don’t know anything about logical decision theory. So I’ll just go on to other questions.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence