Evidence receipt / preference
Published · transcript-backedShane Legg: preference
26 Oct 2023 Dwarkesh Podcast Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures
“We need to make sure it understands humans ethics well, at least as well as a very good ethicist because that's important.”
Source trail
Everything needed to verify it.
- Speaker
- Shane Legg
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 26 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…But if the third one, the ethical model, is a bottleneck, or if it’s a thing that doesn't come along with the AI itself, what is the actual technique to make sure that that happens? First of all, we should train the system on ethics generally so that it understands human ethics well. There's a lot of lectures and papers and books and all sorts of things. We need to make sure it understands humans ethics well, at least as well as a very good ethicist because that's important. And we then need to decide, of this general understanding of ethics, what do we want the system to actually value and what sort of ethics do we want it to apply? Now, that's not a technical problem. That's a problem for society and ethicists and so on to come up with. I'm not sure there's such a thing as optimal ethics but I'm pretty sure that it's possible to come up with a set of ethics, which is much better than what the so-called doomers are worried about in terms of the behavior of these AGI systems. And then what you do is you engineer the system to actually follow these things so that every time it makes a decision, it does an analysis using a deep understanding of the world and of ethics and very robust and precise reasoning to do an ethical analysis of what it's doing. And of course, we would want lots of other things. We would want people checking these processes of reasoning. We’d want people verifying that it's behaving itself in terms of how it reaches these conclusions. But I still feel like I don't understand how that fundamental problem of making sure it follows that ethic works. Because presumably, it has read Mao’s books so it understands Maoist ethics and understands all these other ethics. How do we make sure the ethic that ethicists say is the one is what it ends up following and not the other ones it understands?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.