High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Dwarkesh Patel: uncertainty

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“I’m not trying to kickstart this initial crux again, but I’m just genuinely curious because I think I might be using the term differently.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
uncertainty
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…Humans did it. The researchers did it. Because there’s no other explanation. Gradient descent will not make you generalize well. It will make you solve the problem. It will not make you, if you get new data, generalize in a good way. Generalization means to train on one thing that’ll affect what you do on other things. We know deep learning is really bad at this. For example, we know that if you train on some new thing, it will often catastrophically interfere with all the old things that you knew. This is exactly bad generalization. Generalization, as I said, is some kind of influence of training on one state on other states. The fact that you generalize is not necessarily good or bad. You can generalize poorly, you can generalize well. Generalization always will happen, but we need algorithms that will cause the generalization to be good rather than bad. I’m not trying to kickstart this initial crux again, but I’m just genuinely curious because I think I might be using the term differently. One way to think about these LLMs is that they’re increasing the scope of generalization from earlier systems, which could not really even do a basic math problem, to now where they can do anything in this class of Math Olympiad-type problems. You initially start with them being able to generalize among addition problems. Then they can generalize among problems which require use of different kinds of mathematical techniques and theorems and conceptual categories, which is what the Math Olympiad requires. It sounds like you don’t think of being able to solve any problem within that category as an example of generalization. Let me know if I’m misunderstanding that. Large language models are so complex. We don’t really know what information they have had prior. We have to guess because they’ve been fed so much. This is one reason why they’re not a good way to do science. It’s just so uncontrolled, so unknown.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence