High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Richard Sutton: preference

26 Sept 2025 Dwarkesh Podcast Richard Sutton – Father of RL thinks LLMs are a dead end

“Maybe you’re using “context” because in large language models all that information has to go into the context window.”

— Richard Sutton

Source trail

Everything needed to verify it.

Speaker
Richard Sutton
Attribution
Verified speaker
Claim type
preference
Recorded
26 Sept 2025
Publisher
Dwarkesh Podcast

Transcript context

…It seems to me you need two things. One is some way of converting this long-run goal reward into smaller auxiliary predictive rewards of the future reward, or the future reward that leads to the final reward. But initially, it seems to me, I need to hold on to all this context that I’m gaining as I’m working in the world. I’m learning about my clients, my company, and all this information. I would say you’re just doing regular learning. Maybe you’re using “context” because in large language models all that information has to go into the context window. But in a continual learning setup, it just goes into the weights. Maybe context is the wrong word to use because I mean a more general thing.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence