High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Richard Sutton: preference

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“I think a better word would be “the network” because I think you mean the network.”

— Richard Sutton

Source trail

Everything needed to verify it.

Speaker
Richard Sutton
Attribution
Verified speaker
Claim type
preference
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…I see. With this AI, lots of people would want it to be doing lots of different kinds of things. It’s performing the task people want, but at the same time, it’s learning about the world from doing that task. Let’s say we get rid of this paradigm where there’s training periods and then there’s deployment periods. Do we also get rid of this paradigm where there’s the model and then instances of the model or copies of the model that are doing certain things? How do you think about the fact that we’d want this thing to be doing different things? We’d want to aggregate the knowledge that it’s gaining from doing those different things. I don’t like the word “model” when used the way you just did. I think a better word would be “the network” because I think you mean the network. Maybe there are many networks. Anyway, things would be learned. You’d have copies and many instances. Sure, you’d want to share knowledge across the instances. There would be lots of possibilities for doing that. Today, you have one child grow up and learn about the world, and then every new child has to repeat that process. Whereas with AIs, with a digital intelligence, you could hope to do it once and then copy it into the next one as a starting place. This would be a huge savings. I think it’d be much more important than trying to learn from people. I agree that the kind of thing you’re talking about is necessary regardless of whether you start from LLMs or not. If you want human or animal-level intelligence, you’re going to need this capability. Suppose a human is trying to make a startup. This is a thing which has a reward on the order of 10 years. Once in 10 years you might have an exit where you get paid out a billion dollars. But humans have this ability to make intermediate auxiliary rewards or have some way of…Even when they have extremely sparse rewards, they can still make intermediate steps having an understanding of what the next thing they’re doing leads to this grander goal we have. How do you imagine such a process might play out with AIs?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence