High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“I agree that the kind of thing you’re talking about is necessary regardless of whether you start from LLMs or not.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…I don’t like the word “model” when used the way you just did. I think a better word would be “the network” because I think you mean the network. Maybe there are many networks. Anyway, things would be learned. You’d have copies and many instances. Sure, you’d want to share knowledge across the instances. There would be lots of possibilities for doing that. Today, you have one child grow up and learn about the world, and then every new child has to repeat that process. Whereas with AIs, with a digital intelligence, you could hope to do it once and then copy it into the next one as a starting place. This would be a huge savings. I think it’d be much more important than trying to learn from people. I agree that the kind of thing you’re talking about is necessary regardless of whether you start from LLMs or not. If you want human or animal-level intelligence, you’re going to need this capability. Suppose a human is trying to make a startup. This is a thing which has a reward on the order of 10 years. Once in 10 years you might have an exit where you get paid out a billion dollars. But humans have this ability to make intermediate auxiliary rewards or have some way of…Even when they have extremely sparse rewards, they can still make intermediate steps having an understanding of what the next thing they’re doing leads to this grander goal we have. How do you imagine such a process might play out with AIs? This is something we know very well. The basis of it is temporal difference learning where the same thing happens in a less grandiose scale. When you learn to play chess, you have the long-term goal of winning the game. Yet you want to be able to learn from shorter-term things like taking your opponent’s pieces. You do that by having a value function which predicts the long-term outcome. Then if you take the guy’s pieces, your prediction about the long-term outcome is changed. It goes up, you think you’re going to win. Then that increase in your belief immediately reinforces the move that led to taking the piece. We have this long-term 10-year goal of making a startup and making a lot of money. When we make progress, we say, “Oh, I’m more likely to achieve the long-term goal,” and that rewards the steps along the way.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence