Evidence receipt / evaluation
Published · transcript-backedShane Legg: evaluation
26 Oct 2023 Dwarkesh Podcast Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures
“These are quite big areas. They don't measure things like understanding streaming video, for example, because these are language models and people can do things like understanding streaming video.”
Source trail
Everything needed to verify it.
- Speaker
- Shane Legg
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Let's get more concrete. We measure the performance of these large language models on MMLU and other benchmarks. What is missing from the benchmarks we use currently? What aspect of human cognition do they not measure adequately? Another hard question. These are quite big areas. They don't measure things like understanding streaming video, for example, because these are language models and people can do things like understanding streaming video. They don't do things like episodic memory. Humans have what we call episodic memory. We have a working memory, which are things that have happened quite recently, and then we have a cortical memory, things that are sort of being in our cortex, but there's also a system in between, which is episodic memory, which is the hippocampus. It is about learning specific things very, very rapidly. So if you remember some of the things I say to you tomorrow, that'll be your episodic memory hippocampus. Our models don't really have that kind of thing and we don't really test for that kind of thing. We just sort of try to make the context windows, which is more like working memory, longer and longer to sort of compensate for this. But it is a difficult question because the generality of human intelligence is very, very broad. So you really have to start going into the weeds of trying to find if there's specific types of things that are missing from existing benchmarks or different categories of benchmarks that don't currently exist or something. The thing you're referring to with episodic memory, would it be fair to call that sample efficiency or is that a different thing?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.