Evidence receipt / belief
Published · transcript-backedEliezer Yudkowsky: belief
6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
“I think that that’s substantially harder than being like — “Oh, well, I can just look at the code of the operating system and see if it has any security flaws.”
Source trail
Everything needed to verify it.
- Speaker
- Eliezer Yudkowsky
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 6 Apr 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…So on that second point, I think it would be much easier if both of you had concrete proposals for alignment and you have the pseudocode for alignment. If you’re like “here’s my solution”, and he’s like “here’s my solution.” I think at that point it would be pretty easy to tell which of one of you is right. I think you’re wrong. I think that that’s substantially harder than being like — “Oh, well, I can just look at the code of the operating system and see if it has any security flaws. ” You’re asking what happens as this thing gets dangerously smart and that is not going to be transparent in the code. Let me come back to that. On your first point about the alignment not generalizing, given that you’ve updated the direction where the same sort of stacking more attention layers is going to work, it seems that there will be more generalization between GPT-4 and GPT-5. Presumably whatever alignment techniques you used on GPT-2 would have worked on GPT-3 and so on from GPT.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.