High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Dwarkesh Patel: prediction

26 Oct 2023 Dwarkesh Podcast Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures

“What will it take to align human level and superhuman AIs? It's interesting because the sorts of reinforcement learning and self-play kinds of setups that are popular now, like Constitution AI or RLHF, DeepMind obviously has expertise in it for decades longer.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
prediction
Recorded
26 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…And I can't say much about how we're training. I think it's fair to say we're roughly doing the sorts of scaling and training that you see many people in the field doing but we have our own take on it and our own different tricks and techniques. Okay, maybe we'll come back to it and get another answer on that. Let's talk about alignment briefly. What will it take to align human level and superhuman AIs? It's interesting because the sorts of reinforcement learning and self-play kinds of setups that are popular now, like Constitution AI or RLHF, DeepMind obviously has expertise in it for decades longer. I'm curious what you think of the current landscape and how DeepMind pursues that problem of safety towards human level models. Do you want to know about what we're currently doing or do you want me to have a stab at what I think needs to be done?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence