Evidence receipt / belief
Published · transcript-backedDemis Hassabis: belief
28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold
“I think if a capability like that was discovered through red teaming or external testing, independent testers like government institutes or academia or whatever, then we would have to fix that loophole.”
Source trail
Everything needed to verify it.
- Speaker
- Demis Hassabis
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…When you’re doing these evaluations and for example it turns out your next model could help a layperson build a pandemic-class bioweapon or something, how would you think first of all about making sure those weights are secure so that they don't get out? And second, what would have to be true for you to be comfortable deploying that system? How would you make sure that this latent capability isn’t exposed? The secure model part I think we’ve covered with the cybersecurity and making sure that’s world-class and you’re monitoring all those things. I think if a capability like that was discovered through red teaming or external testing, independent testers like government institutes or academia or whatever, then we would have to fix that loophole. Depending on what it was, that might require a different kind of constitution perhaps, or different guardrails, or more RLHF to avoid that. Or you could remove some training data, depending on what the problem is. I think there could be a number of mitigations. The first part is making sure you detect it ahead of time. So that’s about the right evaluations and right benchmarking and right testing. Then the question is how one would fix that before you deployed it. But I think it would need to be fixed before it was deployed generally, for sure, if that was an exposure surface. Final question. You’ve been thinking in terms of the end goal of AGI at a time when other people thought it was ridiculous in 2010. Now that we’re seeing this slow takeoff where we’re actually seeing generalization and intelligence, what is like psychologically seeing this? What has that been like? Has it just been sort of priced into your world model so it’s not new news for you? Or actually just seeing it live, are you like “wow, something’s really changed”? What does it feel like?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.