High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Richard Sutton: evaluation

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“You can’t have prior knowledge if you don’t have ground truth, because the prior knowledge is supposed to be a hint or an initial belief about what the truth is.”

— Richard Sutton

Source trail

Everything needed to verify it.

Speaker
Richard Sutton
Attribution
Verified speaker
Claim type
evaluation
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…I guess maybe the crux, and I’m curious if you disagree with this, is that some people will say that imitation learning has given us a good prior, or given these models a good prior, of reasonable ways to approach problems. As we move towards the era of experience, as you call it, this prior is going to be the basis on which we teach these models from experience, because this gives them the opportunity to get answers right some of the time. Then on this, you can train them on experience. Do you agree with that perspective? No. I agree that it’s the large language model perspective. I don’t think it’s a good perspective. To be a prior for something, there has to be a real thing. A prior bit of knowledge should be the basis for actual knowledge. What is actual knowledge? There’s no definition of actual knowledge in that large-language framework. What makes an action a good action to take? You recognize the need for continual learning. If you need to learn continually, continually means learning during the normal interaction with the world. There must be some way during the normal interaction to tell what’s right. Is there any way to tell in the large language model setup what’s the right thing to say? You will say something and you will not get feedback about what the right thing to say is, because there’s no definition of what the right thing to say is. There’s no goal. If there’s no goal, then there’s one thing to say, another thing to say. There’s no right thing to say. There’s no ground truth. You can’t have prior knowledge if you don’t have ground truth, because the prior knowledge is supposed to be a hint or an initial belief about what the truth is. There isn’t any truth. There’s no right thing to say. In reinforcement learning, there is a right thing to say, a right thing to do, because the right thing to do is the thing that gets you reward. We have a definition of what’s the right thing to do, so we can have prior knowledge or knowledge provided by people about what the right thing to do is. Then we can check it to see, because we have a definition of what the actual right thing to do is. An even simpler case is when you’re trying to make a model of the world. When you predict what will happen, you predict and then you see what happens. There’s ground truth. There’s no ground truth in large language models because you don’t have a prediction about what will happen next. If you say something in your conversation, the large language models have no prediction about what the person will say in response to that or what the response will be. I think they do. You can literally ask them, “What would you anticipate a user might say in response?” They’ll have a prediction.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence