Evidence receipt / prediction
Published · transcript-backedDwarkesh Patel: prediction
28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold
“Maybe it thinks in alien concepts and you can’t really monitor the million-line pull request because you can’t really understand the whole thing and you can’t give labels.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 28 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…I think if it’s not properly grounded, the system won’t be able to achieve those goals properly. In a sense, you have to have some grounding for a system to actually achieve goals in the real world. I do actually think that these systems, and things like Gemini, are becoming more multimodal. As we start ingesting things like video and audiovisual data as well as text data, then the system starts correlating those things together. I think that is a form of proper grounding. So I do think our systems are going to start to understand the physics of the real world better. Then one could imagine the active version of that as a very realistic simulation or game environment where you’re starting to learn about what your actions do in the world and how that affects the world itself. The world stays itself, but it also affects what next learning episode you’re getting. So these RL agents we’ve always been working on and pioneered, like AlphaZero and AlphaGo, actually are active learners. What they decide to do next affects what next learning piece of data or experience they’re going to get. So there’s this very interesting sort of feedback loop. And of course, if we ever want to be good at things like robotics, we’re going to have to understand how to act in the real world. So there’s grounding in terms of whether the capabilities will be able to proceed, whether they will be enough in touch with reality to do the things we want. There’s another sense of grounding in that we’ve gotten lucky that since they’re trained on human thought, they maybe think like a human. To what extent does that stay true when more of the compute for training comes from just “did you get the right outcome” and it’s not guardrailed by “are you proceeding on the next token as a human would?” Maybe the broader question I’ll pose to you is, and this is what I asked Shane as well, what would it take to align a system that’s smarter than a human? Maybe it thinks in alien concepts and you can’t really monitor the million-line pull request because you can’t really understand the whole thing and you can’t give labels. This is something Shane and I, and many others here, have had at the forefront of our minds since before we started DeepMind because we planned for success. In 2010, no one was thinking about AI let alone AGI. But we already knew that if we could make progress with these systems and these ideas, the technology created would be unbelievably transformative. So we were already thinking 20 years ago about what the consequences of that would be, both positive and negative. Of course, the positive direction is amazing science, things like AlphaFold, incredible breakthroughs in health and science, and mathematical and scientific discovery. But we also have to make sure these systems are sort of understandable and controllable. This will be a whole discussion in itself, but there are many, many ideas that people have such as more stringent eval systems. I think we don’t have good enough evaluations and benchmarks for things like if the system can deceive you. Can it exfiltrate its own code or do other undesirable behaviors? There are also ideas of using AI, not general learning ones but maybe narrow AIs that are specialized for a domain, to help us as the human scientists to analyze and summarize what the more general system is doing. So there’s narrow AI tools. I think that there’s a lot of promise in creating hardened sandboxes or simulations that are hardened with cybersecurity arrangements around the simulation, both to keep the AI in and to keep hackers out. You could experiment a lot more freely within that sandbox domain. There’s many, many other ideas, including the analysis stuff we talked about earlier, where we can analyze and understand what the concepts are that this system is building and what the representations are. So maybe then they’re not so alien to us and we can actually keep track of the kind of knowledge that it’s building.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.