Evidence receipt / evaluation
Published · transcript-backedDemis Hassabis: evaluation
28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold
“I think we get some grounding through the RLHF feedback systems because obviously the human raters are by definition, grounded people.”
Source trail
Everything needed to verify it.
- Speaker
- Demis Hassabis
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 28 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…This sounds like the answer to the question I’m about to ask. As somebody who’s been in this field for a long time and seen different trends come and go, what do you think the strong version of the scaling hypothesis gets right and what does it get wrong? The idea that you just throw enough compute at a wide enough distribution of data and you get intelligence. My view is that this is kind of an empirical question right now. I think it was pretty surprising to almost everyone, including the people who first worked on the scaling hypotheses, how far it’s gone. In a way, I look at the large models today and I think they’re almost unreasonably effective for what they are. I think it’s pretty surprising some of the properties that emerge. In my opinion, they’ve clearly got some form of concepts and abstractions and things like that. I think if we were talking five-plus years ago, I would have said to you that maybe we need an additional algorithmic breakthrough in order to do that, maybe more like how the brain works. I think that’s still true if we want explicit abstract concepts, neat concepts, but it seems that these systems can implicitly learn that. Another really interesting, unexpected thing was that these systems have some sort of grounding even though they don’t experience the world multimodally, at least until more recently when we have the multimodal models. The amount of information and models that can be built up just from language is surprising. I think that I’d have some hypotheses about why that is. I think we get some grounding through the RLHF feedback systems because obviously the human raters are by definition, grounded people. We’re grounded in reality, so our feedback is also grounded. Perhaps there’s some grounding coming in through there. Also if you’re able to ingest all of it, maybe language contains more grounding than linguists thought before. So it actually raises some very interesting philosophical questions that people haven’t even really scratched the surface of yet. Looking at the advances that have been made, it’s quite interesting to think about where it’s going to go next. In terms of your question of large models, I think we’ve got to push scaling as hard as we can and that’s what we’re doing here. It’s an empirical question, whether that will hit an asymptote or a brick wall, and there are different people who argue about that. I think we should just test it. I think no one knows. In the meantime, we should also double down on innovation and invention. This is something where Google Research and DeepMind and Google Brain have pioneered many, many things over the last decade. That’s our bread and butter. You can think of half our effort as having to do with scaling and half our efforts having to do with inventing the next architectures and the next algorithms that will be needed, knowing that larger and larger scaled models are coming down the line. So my betting right now, but it’s a loose betting, is that you need both. I think you’ve got to push both of them as hard as possible and we’re in a lucky position that we can do that. I want to ask more about the grounding. You can imagine two things that might change which would make the grounding more difficult. One is that as these models get smarter, they are going to be able to operate in domains where we just can’t generate enough human labels, just because we’re not smart enough. If it does a million-line pull request, how do we tell it, for example, this is within the constraints of our morality and the end goal we wanted and this isn’t? The other thing has to do with what you were saying about compute. So far we’ve been doing next token prediction and in some sense it’s a guardrail, because you have to talk as a human would talk and think as a human would think. Now, additional compute is maybe going to come in the form of reinforcement learning where it’s just getting to the objective and we can’t really trace how you got there. When you combine those two, how worried are you that the grounding goes away?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.