Evidence receipt / preference
Published · transcript-backedRichard Sutton: preference
26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end
“Maybe you’re using “context” because in large language models all that information has to go into the context window.”
Source trail
Everything needed to verify it.
- Speaker
- Richard Sutton
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 26 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…It seems to me you need two things. One is some way of converting this long-run goal reward into smaller auxiliary predictive rewards of the future reward, or the future reward that leads to the final reward. But initially, it seems to me, I need to hold on to all this context that I’m gaining as I’m working in the world. I’m learning about my clients, my company, and all this information. I would say you’re just doing regular learning. Maybe you’re using “context” because in large language models all that information has to go into the context window. But in a continual learning setup, it just goes into the weights. Maybe context is the wrong word to use because I mean a more general thing.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.