Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end
“I’m not trying to kickstart this initial crux again, but I’m just genuinely curious because I think I might be using the term differently.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 26 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Humans did it. The researchers did it. Because there’s no other explanation. Gradient descent will not make you generalize well. It will make you solve the problem. It will not make you, if you get new data, generalize in a good way. Generalization means to train on one thing that’ll affect what you do on other things. We know deep learning is really bad at this. For example, we know that if you train on some new thing, it will often catastrophically interfere with all the old things that you knew. This is exactly bad generalization. Generalization, as I said, is some kind of influence of training on one state on other states. The fact that you generalize is not necessarily good or bad. You can generalize poorly, you can generalize well. Generalization always will happen, but we need algorithms that will cause the generalization to be good rather than bad. I’m not trying to kickstart this initial crux again, but I’m just genuinely curious because I think I might be using the term differently. One way to think about these LLMs is that they’re increasing the scope of generalization from earlier systems, which could not really even do a basic math problem, to now where they can do anything in this class of Math Olympiad-type problems. You initially start with them being able to generalize among addition problems. Then they can generalize among problems which require use of different kinds of mathematical techniques and theorems and conceptual categories, which is what the Math Olympiad requires. It sounds like you don’t think of being able to solve any problem within that category as an example of generalization. Let me know if I’m misunderstanding that. Large language models are so complex. We don’t really know what information they have had prior. We have to guess because they’ve been fed so much. This is one reason why they’re not a good way to do science. It’s just so uncontrolled, so unknown.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.