Evidence receipt / evaluation
Published · transcript-backedRichard Sutton: evaluation
26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end
“Your examples are all, “Well, really you have to” because you can teach it, but there’s all the little idiosyncrasies of the particular life they’re leading and the particular people they’re working with and what they like, as opposed to what average people like.”
Source trail
Everything needed to verify it.
- Speaker
- Richard Sutton
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…You also want some ability for information that you’re learning. One of the things that makes humans quite different from these LLMs is that if you’re onboarding on a job, you’re picking up so much context and information. That’s what makes you useful at the job. You’re learning everything from how your client has preferences to how the company works, everything. Is the bandwidth of information that you get from a procedure like TD learning high enough to have this huge pipe of context and tacit knowledge that you need to be picking up in the way humans do when they’re just deployed? I’m not sure but I think at the crux of this, the big world hypothesis seems very relevant. The reason why humans become useful on the job is because they are encountering their particular part of the world. It can’t have been anticipated and can’t all have been put in in advance. The world is so huge that you can’t. The dream of large language models, as I see it, is you can teach the agent everything. It will know everything and won’t have to learn anything online, during its life. Your examples are all, “Well, really you have to” because you can teach it, but there’s all the little idiosyncrasies of the particular life they’re leading and the particular people they’re working with and what they like, as opposed to what average people like. That’s just saying the world is really big, and you’re going to have to learn it along the way. It seems to me you need two things. One is some way of converting this long-run goal reward into smaller auxiliary predictive rewards of the future reward, or the future reward that leads to the final reward. But initially, it seems to me, I need to hold on to all this context that I’m gaining as I’m working in the world. I’m learning about my clients, my company, and all this information.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.