Evidence receipt / evaluation
Published · transcript-backedSholto Douglas: evaluation
25 Mar 2025 Dwarkesh Podcast AMA: career advice given AGI, how I research ft. Sholto & Trenton
“I think my answer at the moment is that the sort of pre-training objective doesn't necessarily- like it imbues with this nice flexible general knowledge about the world, but doesn't necessarily imbue the skill of making novel connections or research.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Mar 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…So the issue is, one of the questions I asked Dario is, look, these models have all of human knowledge memorized and you would think if a human had this much stuff memorized, and they were moderately intelligent, they could be making all these connections between different fields. And there are examples of humans doing this, by the way. There's… Donald Swann or something like this [Don R. Swanson], this guy noticed that what happens to a brain after magnesium deficiency is exactly the structure you see during a migraine. So then he's like, you take magnesium supplements and we're gonna cure a bunch of migraines. And it worked. And there's many other examples of things like this where you just notice two different connections between pieces of knowledge. Why, if these LLMs are intelligent, are they not able to use this unique advantage they have to make these kinds of discoveries? I feel a little shy, me giving answers on AI shit with you guys here. But, actually Scott Alexander addressed this question in one of his AMA threads, and he's like, "Look, humans also don't have this kind of logical omniscience", right? He used the example of, in language, if you really thought about, why are two words connected? And it's like, I understand why “rhyme” has the same etymology as this other word. But you just don't think about it, right? There's this combinatorial explosion. I don't know if that addresses the fact that- we know humans can do this, right? The humans have in fact done this, and I don't know of a single example of LLMs ever having done it. Actually, yeah, what is your answer to this? I think my answer at the moment is that the sort of pre-training objective doesn't necessarily- like it imbues with this nice flexible general knowledge about the world, but doesn't necessarily imbue the skill of making novel connections or research. The kinds of things that people are trained to do through PhD programs and through the process of exploring and interacting with the world. And so I think at a minimum you need significant RL in at least similar things to be able to approach making novel discoveries. And so I would like to see some early evidence of this as we start to build models that are interacting with the world and trying to make scientific discoveries, and modeling the behaviors that we expect of people in these positions. Because I don't actually think we've done that in a meaningful or scaled way as a field, so to speak. Riffing off that with respect to RL, I wonder if models currently just aren't good at knowing what memories they should be storing. Most of their training is just predicting the next word on the internet and remembering very specific facts from that. But if you were to teach me something new right now, I'm very aware of my own memory limitations, and so I would try to construct some summary that would stick. And models currently don't have the opportunity to do that. Memory scaffolding in general is just very primitive right now. I mean-…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.