Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
“The way you describe it, it seemed kind of compelling. I don’t know why that doesn’t even rise to 1%.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 6 Apr 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…nding to be. I can say this and the disaster monkeys at the current places cannot (unclear) to it but they have not said things like this themselves that I have ever heard and that is not a good sign. And then if you don’t amp this up too far, which on the present paradigm you can’t do anyways, because if you train the very, very smart version of the system it kills you before you can RLHF it. But maybe you can train GPT to distinguish nice, valid, kind, careful, and then filter all the training data to get the nice things to train on and then train on that data rather than training on everything to try to avert the Waluigi problem or just more generally having all the darkness in there. Just train it on the light that’s in humanity. So there’s like that kind of course. And if you don’t push that too far, maybe you can get a genuine ally and maybe things play out differently from there. That’s one of the little rays of hope. But I don’t think alignment is actually so easy that you just get whatever you want. It’s a genie, it gives you what you wish for. I don’t think that doesn’t even strike me as hope. Honestly. The way you describe it, it seemed kind of compelling. I don’t know why that doesn’t even rise to 1%. The possibility that it works out that way. This is like literally my AI alignment fantasy from 2003, though not with RLHF as the implementation method or LLMs as the base. And it’s going to be more dangerous than what I was dreaming about in 2003. And I think in a very real sense it feels to me like the people doing this stuff now have literally not gotten as far as I was in 2003. And I’ve now written out my answer sheet for that. It’s on the podcast, it goes on the Internet. And now they can pretend that that was their idea or like — “Sure, that’s obvious. We were going to do that anyways.” And yet they didn’t say it earlier. You can’t run a big project off of one person who.. The alignment field failed to gel. That’s my (unclear) to the like — “Well, you just throw in a ton of more money, and then it’s all solvable.” Because I’ve seen people try to amp up the amount of money that goes into it and the stuff coming out of it has not gone to the places that I would have considered obvious a while ago. And I can print out all my answer sheets for it and each time I do that, it gets a little bit harder to make the case next time.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.