Evidence receipt / belief
Published · transcript-backedDemis Hassabis: belief
28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold
“I think what we’ve got to do in the next few years, in the time before those systems start arriving, is come up with the right evaluations and metrics.”
Source trail
Everything needed to verify it.
- Speaker
- Demis Hassabis
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…I’m curious what you think about that. I’m not saying this is happening this year, but eventually you’ll be developing a model where you think there’s some chance that it’ll be capable of an intelligence explosion-like dynamic once it’s fully developed. What would have to be true of that model at that point where you’re comfortable continuing the development of the system? Something like, “I’ve seen these specific evals, I’ve understood its internal thinking and its future thinking enough.” We need a lot more understanding of the systems than we do today before I would even be confident of explaining to you what we’d need to tick box there. I think what we’ve got to do in the next few years, in the time before those systems start arriving, is come up with the right evaluations and metrics. Ideally formal proofs, but it’s going to be hard for these types of systems, so at least empirical bounds around what these systems can do. That’s why I think about things like deception as being quite root node traits that you don’t want. If you’re confident that your system is exposing what it actually thinks, then that opens up possibilities of using the system itself to explain aspects of itself to you. The way I think about that is like this. If I were to play a game of chess against Garry Kasparov, which I’ve played in the past, Magnus Carlsen, or the amazing chess players of all time, I wouldn’t be able to come up with a move that they could. But they could explain to me why they came up with that move and I could understand it post hoc, right? That’s the sort of thing one could imagine. One of the capabilities that we could make use of these systems is for them to explain it to us and even maybe get the proofs behind why they’re thinking something, certainly in a mathematical problem. Got it. Do you have a sense of what the converse answer would be? So what would have to be true where tomorrow morning you’re like “oh, man, I didn’t anticipate this.” You see some specific observation tomorrow morning that makes you say “we got to stop Gemini 2 training.”…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.