Evidence receipt / prediction
Published · transcript-backedShane Legg: prediction
26 Oct 2023 Dwarkesh Podcast Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures
“I think the next landmark that people will think back to and remember is going much more fully multimodal.”
Source trail
Everything needed to verify it.
- Speaker
- Shane Legg
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 26 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…The final question is this. You've been in this field for over a decade, much longer than many others, and you've seen different landmarks like ImageNet and Transformers. What do you think the next landmark will look like? I think the next landmark that people will think back to and remember is going much more fully multimodal. That will open out the sort of understanding that you see in language models into a much larger space of possibilities. And when people think back, they'll think about, “Oh, those old fashioned models, they just did like chat, they just did text.” It just felt like a very narrow thing whereas now they understand when you talk to them and they understand images and pictures and video and you can show them things or things like that. And they will have much more understanding of what's going on. And it'll feel like the system's kind of opened up into the world in a much more powerful way. Do you mind if I ask a follow-up on that? ChatGPT just released their multimodal feature and you, in DeepMind, you had the Gato paper, where you have this one model where you can throw images, video games and even actions in there. So far it doesn't seem to have percolated as much as ChatGPT initially from GPT3 or something. What explains that? Is it just that people haven't learned to use multimodality? They're not powerful enough yet?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.