Evidence receipt / prediction
Published · transcript-backedDwarkesh Patel: prediction
28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold
“One is that as these models get smarter, they are going to be able to operate in domains where we just can’t generate enough human labels, just because we’re not smart enough.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 28 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…My view is that this is kind of an empirical question right now. I think it was pretty surprising to almost everyone, including the people who first worked on the scaling hypotheses, how far it’s gone. In a way, I look at the large models today and I think they’re almost unreasonably effective for what they are. I think it’s pretty surprising some of the properties that emerge. In my opinion, they’ve clearly got some form of concepts and abstractions and things like that. I think if we were talking five-plus years ago, I would have said to you that maybe we need an additional algorithmic breakthrough in order to do that, maybe more like how the brain works. I think that’s still true if we want explicit abstract concepts, neat concepts, but it seems that these systems can implicitly learn that. Another really interesting, unexpected thing was that these systems have some sort of grounding even though they don’t experience the world multimodally, at least until more recently when we have the multimodal models. The amount of information and models that can be built up just from language is surprising. I think that I’d have some hypotheses about why that is. I think we get some grounding through the RLHF feedback systems because obviously the human raters are by definition, grounded people. We’re grounded in reality, so our feedback is also grounded. Perhaps there’s some grounding coming in through there. Also if you’re able to ingest all of it, maybe language contains more grounding than linguists thought before. So it actually raises some very interesting philosophical questions that people haven’t even really scratched the surface of yet. Looking at the advances that have been made, it’s quite interesting to think about where it’s going to go next. In terms of your question of large models, I think we’ve got to push scaling as hard as we can and that’s what we’re doing here. It’s an empirical question, whether that will hit an asymptote or a brick wall, and there are different people who argue about that. I think we should just test it. I think no one knows. In the meantime, we should also double down on innovation and invention. This is something where Google Research and DeepMind and Google Brain have pioneered many, many things over the last decade. That’s our bread and butter. You can think of half our effort as having to do with scaling and half our efforts having to do with inventing the next architectures and the next algorithms that will be needed, knowing that larger and larger scaled models are coming down the line. So my betting right now, but it’s a loose betting, is that you need both. I think you’ve got to push both of them as hard as possible and we’re in a lucky position that we can do that. I want to ask more about the grounding. You can imagine two things that might change which would make the grounding more difficult. One is that as these models get smarter, they are going to be able to operate in domains where we just can’t generate enough human labels, just because we’re not smart enough. If it does a million-line pull request, how do we tell it, for example, this is within the constraints of our morality and the end goal we wanted and this isn’t? The other thing has to do with what you were saying about compute. So far we’ve been doing next token prediction and in some sense it’s a guardrail, because you have to talk as a human would talk and think as a human would think. Now, additional compute is maybe going to come in the form of reinforcement learning where it’s just getting to the objective and we can’t really trace how you got there. When you combine those two, how worried are you that the grounding goes away? I think if it’s not properly grounded, the system won’t be able to achieve those goals properly. In a sense, you have to have some grounding for a system to actually achieve goals in the real world. I do actually think that these systems, and things like Gemini, are becoming more multimodal. As we start ingesting things like video and audiovisual data as well as text data, then the system starts correlating those things together. I think that is a form of proper grounding. So I do think our systems are going to start to understand the physics of the real world better. Then one could imagine the active version of that as a very realistic simulation or game environment where you’re starting to learn about what your actions do in the world and how that affects the world itself. The world stays itself, but it also affects what next learning episode you’re getting. So these RL agents we’ve always been working on and pioneered, like AlphaZero and AlphaGo, actually are active learners. What they decide to do next affects what next learning piece of data or experience they’re going to get. So there’s this very interesting sort of feedback loop. And of course, if we ever want to be good at things like robotics, we’re going to have to understand how to act in the real world.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.