High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Dwarkesh Patel: observation

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“It’ll say, “Okay, I’m going to approach this problem using this approach first.” It’ll write this out and be like, “Oh wait, I just realized this is the wrong conceptual way to approach the problem.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
observation
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…No, they will respond to that question right. But they have no prediction in the substantive sense that they won’t be surprised by what happens. If something happens that isn’t what you might say they predicted, they will not change because an unexpected thing has happened. To learn that, they’d have to make an adjustment. I think a capability like this does exist in context. It’s interesting to watch a model do chain of thought. Suppose it’s trying to solve a math problem. It’ll say, “Okay, I’m going to approach this problem using this approach first.” It’ll write this out and be like, “Oh wait, I just realized this is the wrong conceptual way to approach the problem. I’m going to restart with another approach.” That flexibility does exist in context, right? Do you have something else in mind or do you just think that you need to extend this capability across longer horizons? I’m just saying they don’t have in any meaningful sense a prediction of what will happen next. They will not be surprised by what happens next. They’ll not make any changes if something happens, based on what happens.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence