Evidence receipt / prediction
Published · transcript-backedDwarkesh Patel: prediction
26 Oct 2023 Dwarkesh Podcast Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures
“What will it take to align human level and superhuman AIs? It's interesting because the sorts of reinforcement learning and self-play kinds of setups that are popular now, like Constitution AI or RLHF, DeepMind obviously has expertise in it for decades longer.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 26 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…And I can't say much about how we're training. I think it's fair to say we're roughly doing the sorts of scaling and training that you see many people in the field doing but we have our own take on it and our own different tricks and techniques. Okay, maybe we'll come back to it and get another answer on that. Let's talk about alignment briefly. What will it take to align human level and superhuman AIs? It's interesting because the sorts of reinforcement learning and self-play kinds of setups that are popular now, like Constitution AI or RLHF, DeepMind obviously has expertise in it for decades longer. I'm curious what you think of the current landscape and how DeepMind pursues that problem of safety towards human level models. Do you want to know about what we're currently doing or do you want me to have a stab at what I think needs to be done?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.