Evidence receipt / belief
Published · transcript-backedSholto Douglas: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“The residual stream is like operating RAM, you're doing stuff to it, is the mental model I think one takes away from interpretability work.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…I mean, you can already think of models... Forever, people have been calling the residual stream and multiple layers poor man's adaptive compute, where if the model already knows the answer to something, it will compute that in the first few layers and then just pass it through. I mean, that's getting into the weeds. The residual stream is like operating RAM, you're doing stuff to it, is the mental model I think one takes away from interpretability work. We've been talking a lot about scratchpads, them writing down their thoughts in ways in which they're already unreliable in some respects. Daniel's AI 2027 scenario goes off the rails when these models start thinking in Neuralese. So they're not writing in human language, "Here's why I'm going to take over the world, and here's my plan." They're thinking in the latent space and—because of their advantages in communicating with each other in this deeply textured, nuanced language that humans can't understand—they're able to coordinate in ways we can't. Is this the path for future models? Are they going to be, in Neuralese, communicating with themselves or with each other?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.