Evidence receipt / recommendation
Published · transcript-backedEliezer Yudkowsky: recommendation
6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
“I could say go study evolutionary biology because evolutionary biology went through a phase of optimism and people naming all the wonderful things they thought that evolutionary biology would cough out, all the wonderful properties that they thought natural selection would imbue into organisms.”
Source trail
Everything needed to verify it.
- Speaker
- Eliezer Yudkowsky
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 6 Apr 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Okay, final question. I don’t know how many hours this has been. I really appreciate you giving me your time. I know that in a previous episode, you were not able to give specific advice of what somebody young who is motivated to work on these problems should do. Do you have advice about how one would even approach coming up with an answer to that themselves? There’s people running programs who think we have more time, who think we have better chances, and they’re running programs to try to nudge people into doing useful work in this area. And I’m not sure they’re working. And there’s such a strange road to walk and not a short one. And I tried to help people along the way, and I don’t think they got far enough. Some of them got some distance, but they didn’t turn into alignment specialists doing great work. And it’s the problem of the broken verifier. If somebody had a bunch of talent in physics, they were like — Well, I want to work in this field. I might be like — Well, there’s interpretability, and you can tell whether you’ve made a discovery in interpretability or not. Sets it apart for a bunch of this other stuff, and I don’t think that saves us. So how do you do the kind of work that saves us? The key thing is the ability to tell the difference between good and bad work. And maybe I will write some more blog posts on it. I don’t really expect the blog posts to work. The critical thing is the verifier. How can you tell whether you’re talking sense or not? There’s all kinds of specific heuristics I can give. I can say to somebody — “Well, if your entire alignment proposal is this elaborate mechanism you have to explain the whole mechanism.” And you can’t be like “here’s the core problem. Here’s the key insight that I think addresses this problem.” If you can’t extract that out, if your whole solution is just a giant mechanism, this is not the way. It’s kind of like how people invent perpetual motion machines by making the perpetual motion machines more and more complicated until they can no longer keep track of how it fails. And if you actually had a perpetual motion machine, it would not just be a giant machine, there would be a thing you had realized that made it possible to do the impossible, for example. You’re just not going to have a perpetual motion machine. So there’s thoughts like that. I could say go study evolutionary biology because evolutionary biology went through a phase of optimism and people naming all the wonderful things they thought that evolutionary biology would cough out, all the wonderful properties that they thought natural selection would imbue into organisms. And the Williams Revolution as is sometimes called, is when George Williams wrote Adaptation and Natural Selection, a very influential book. Saying like that is not what this optimization criterion gives you. You do not get the pretty stuff, you do not get the aesthetically lovely stuff. Here’s what you get instead. And by living through that revolution vicariously. I thereby picked up a bit of the thing that to me obviously generalizes about how not to expect nice things from an alien optimization process. But maybe somebody else can read through that and not generalize in the correct direction. So then how do I advise them to generalize in the correct direction? ings from an alien optimization process. But maybe somebody else can read through that and not generalize in the correct direction. So then how do I advise them to generalize in the correct direction? How do I advise them to learn the thing that I learned? I can just give them the generalization but that’s not the same as having the thing inside them that generalizes correctly without anybody standing over their shoulder and forcing them to get the right answer. I could point out and have in my fiction that the entire schooling process of — “Here is this legible question that you’re supposed to have already been taught how to solve. Give me the answer using the solution method you are taught.” This does not train you to tackle new basic problems. But even if you tell people that, how do they retrain? We don’t have a systematic training method for producing real science in that sense. A quarter of the Nobel laureates being the students or grad students of other Nobel laureates because we never figured out how to teach science. We have an apprentice system. We have people who pick out people who they think can be scientists and they hang around them in person. And something that we’ve never written down in a textbook passes down. And that’s where the revolutionaries come from. And there are whole countries trying to invest in having scientists, and they churn out these people who write papers, and none of it goes anywhere. Because the part that was legible to the bureaucracy is, have you written the paper? Can you pass the test? And this is not science. And I could go on for this for a while, but the thing that you asked me is — How do you pass down this thing that your society never did figure out how to teach? And the whole reason why Harry Potter and the Methods of Rationality is popular is because people read it and picked up the rhythm seen in a character’s thoughts of a thing that was not in their schooling system, that was not written down, that you would ordinarily pick up by being around other people. And I managed to put a little bit of it into a fictional character, and people picked up a fragment of it by being near a fictional character, but not in really vast quantities of people. And I didn’t manage to put vast quantities of shards in there. I’m not sure there is not a long list of Nobel laureates who’ve read HPMOR, although there wouldn’t be, because the delay times on granting the prizes are too long. You ask me, what do I say? And my answer is — Well, that’s a whole big, gigantic problem I’ve spent however many years trying to tackle, and I ain’t going to solve the problem with a sentence in this podcast.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.