High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Demis Hassabis: belief

28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold

“I think if it’s not properly grounded, the system won’t be able to achieve those goals properly.”

— Demis Hassabis

Source trail

Everything needed to verify it.

Speaker
Demis Hassabis
Attribution
Verified speaker
Claim type
belief
Recorded
28 Feb 2024
Publisher
Dwarkesh Podcast

Transcript context

…I want to ask more about the grounding. You can imagine two things that might change which would make the grounding more difficult. One is that as these models get smarter, they are going to be able to operate in domains where we just can’t generate enough human labels, just because we’re not smart enough. If it does a million-line pull request, how do we tell it, for example, this is within the constraints of our morality and the end goal we wanted and this isn’t? The other thing has to do with what you were saying about compute. So far we’ve been doing next token prediction and in some sense it’s a guardrail, because you have to talk as a human would talk and think as a human would think. Now, additional compute is maybe going to come in the form of reinforcement learning where it’s just getting to the objective and we can’t really trace how you got there. When you combine those two, how worried are you that the grounding goes away? I think if it’s not properly grounded, the system won’t be able to achieve those goals properly. In a sense, you have to have some grounding for a system to actually achieve goals in the real world. I do actually think that these systems, and things like Gemini, are becoming more multimodal. As we start ingesting things like video and audiovisual data as well as text data, then the system starts correlating those things together. I think that is a form of proper grounding. So I do think our systems are going to start to understand the physics of the real world better. Then one could imagine the active version of that as a very realistic simulation or game environment where you’re starting to learn about what your actions do in the world and how that affects the world itself. The world stays itself, but it also affects what next learning episode you’re getting. So these RL agents we’ve always been working on and pioneered, like AlphaZero and AlphaGo, actually are active learners. What they decide to do next affects what next learning piece of data or experience they’re going to get. So there’s this very interesting sort of feedback loop. And of course, if we ever want to be good at things like robotics, we’re going to have to understand how to act in the real world. So there’s grounding in terms of whether the capabilities will be able to proceed, whether they will be enough in touch with reality to do the things we want. There’s another sense of grounding in that we’ve gotten lucky that since they’re trained on human thought, they maybe think like a human. To what extent does that stay true when more of the compute for training comes from just “did you get the right outcome” and it’s not guardrailed by “are you proceeding on the next token as a human would?” Maybe the broader question I’ll pose to you is, and this is what I asked Shane as well, what would it take to align a system that’s smarter than a human? Maybe it thinks in alien concepts and you can’t really monitor the million-line pull request because you can’t really understand the whole thing and you can’t give labels.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence