Evidence receipt / belief
Published · transcript-backedRichard Sutton: belief
26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end
“I think there also should be a component having to do with your increasing understanding of your environment.”
Source trail
Everything needed to verify it.
- Speaker
- Richard Sutton
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Sept 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, I guess what I meant to say is a general human-level, general continual learning agent. What is the reward function? Is it just predicting the world? Is it then having a specific effect on it? What would the general reward function be? The reward function is arbitrary. If you’re playing chess, it’s to win the game of chess. If you’re a squirrel, maybe the reward has to do with getting nuts. In general, for an animal, you would say the reward is to avoid pain and to acquire pleasure. I think there also should be a component having to do with your increasing understanding of your environment. That would be sort of an intrinsic motivation. I see. With this AI, lots of people would want it to be doing lots of different kinds of things. It’s performing the task people want, but at the same time, it’s learning about the world from doing that task. Let’s say we get rid of this paradigm where there’s training periods and then there’s deployment periods. Do we also get rid of this paradigm where there’s the model and then instances of the model or copies of the model that are doing certain things? How do you think about the fact that we’d want this thing to be doing different things? We’d want to aggregate the knowledge that it’s gaining from doing those different things.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.