Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end
“It’s interesting because you wrote this essay in 2019 titled “The Bitter Lesson,” and this is the most influential essay, perhaps, in the history of AI. But people have used that as a justification for scaling up LLMs because, in their view, this is the one scalable way we have found to pour ungodly amounts of compute into learning about the world.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…The math problems are different. Making a model of the physical world and carrying out the consequences of mathematical assumptions or operations, those are very different things. The empirical world has to be learned. You have to learn the consequences. Whereas the math is more computational, it’s more like standard planning. There they can have a goal to find the proof, and they are in some way given that goal to find the proof. It’s interesting because you wrote this essay in 2019 titled “The Bitter Lesson,” and this is the most influential essay, perhaps, in the history of AI. But people have used that as a justification for scaling up LLMs because, in their view, this is the one scalable way we have found to pour ungodly amounts of compute into learning about the world. It’s interesting that your perspective is that the LLMs are not “bitter lesson”-pilled. It’s an interesting question whether large language models are a case of the bitter lesson. They are clearly a way of using massive computation, things that will scale with computation up to the limits of the Internet. But they’re also a way of putting in lots of human knowledge. This is an interesting question. It’s a sociological or industry question. Will they reach the limits of the data and be superseded by things that can get more data just from experience rather than from people? In some ways it’s a classic case of the bitter lesson. The more human knowledge we put into the large language models, the better they can do. So it feels good. Yet, I expect there to be systems that can learn from experience. Which could perform much better and be much more scalable. In which case, it will be another instance of the bitter lesson, that the things that used human knowledge were eventually superseded by things that just trained from experience and computation.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.