High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Richard Sutton: belief

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“I think there also should be a component having to do with your increasing understanding of your environment.”

— Richard Sutton

Source trail

Everything needed to verify it.

Speaker
Richard Sutton
Attribution
Verified speaker
Claim type
belief
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…Yeah, I guess what I meant to say is a general human-level, general continual learning agent. What is the reward function? Is it just predicting the world? Is it then having a specific effect on it? What would the general reward function be? The reward function is arbitrary. If you’re playing chess, it’s to win the game of chess. If you’re a squirrel, maybe the reward has to do with getting nuts. In general, for an animal, you would say the reward is to avoid pain and to acquire pleasure. I think there also should be a component having to do with your increasing understanding of your environment. That would be sort of an intrinsic motivation. I see. With this AI, lots of people would want it to be doing lots of different kinds of things. It’s performing the task people want, but at the same time, it’s learning about the world from doing that task. Let’s say we get rid of this paradigm where there’s training periods and then there’s deployment periods. Do we also get rid of this paradigm where there’s the model and then instances of the model or copies of the model that are doing certain things? How do you think about the fact that we’d want this thing to be doing different things? We’d want to aggregate the knowledge that it’s gaining from doing those different things.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence