High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eliezer Yudkowsky: evaluation

6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

“The academic literature would have to be seen to be believed. But the point is the one major technical contribution that I’m proud of, which is not all that precedented and you can look at the literature and see it’s not all that precedented, would in fact have been a way for something that knew about that technical innovation to build a superintelligence that would kill you and extract value itself from that superintelligence in a way that would just completely blindside the literature as it existed prior to that technical contribution.”

— Eliezer Yudkowsky

Source trail

Everything needed to verify it.

Speaker
Eliezer Yudkowsky
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 Apr 2023
Publisher
Dwarkesh Podcast

Transcript context

…On the object level, I don’t know whether somebody could have figured that out because I’m not sure what the thing is. The academic literature would have to be seen to be believed. But the point is the one major technical contribution that I’m proud of, which is not all that precedented and you can look at the literature and see it’s not all that precedented, would in fact have been a way for something that knew about that technical innovation to build a superintelligence that would kill you and extract value itself from that superintelligence in a way that would just completely blindside the literature as it existed prior to that technical contribution. And there’s going to be other stuff like that. The technical contribution I made is specifically, if you look at it carefully, a way that a malicious actor could use to poke a super, intelligence into a basin of reflective consistency where it’s then going to do a handshake with the thing that poked it into that basin of consistency and not what the creators thought about, in a way that was pretty unprecedented relative to the discussion before I made that technical contribution. Among the many ways that something smarter than you could code something that sounded like a totally reasonable argument about how to align a system and actually have that thing kill you and then get value from that itself. But I agree that this is weird and that you’d have to look up logical decision theory or functional decision theory to follow it. Yeah, I can’t evaluate that at an object level right now.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence