Evidence receipt / observation
Published · transcript-backedSholto Douglas: observation
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“Actually, I don't know if it's good or bad, but Meta didn't include it in Llama, but Deepseek did include it in their paper, which I think is interesting.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Does all training involve continuously fucking with these values as you're going through it? It depends on what your architecture is. But I thought it was just cute that you can see them running up into this very hardware-level constraint, trying to go, "What do we wish we could express algorithmically? What can we express under our constraints?" and iteratively solving to get better constraints. And doing this in a really simple and elegant way, and then backing it up with great engineering. I also thought it was interesting that they incorporated the multi-token prediction thing from Meta. So Meta had a nice paper on this multi-token prediction thing. Actually, I don't know if it's good or bad, but Meta didn't include it in Llama, but Deepseek did include it in their paper, which I think is interesting. Was that because they were faster at iterating and including an algorithm? Or did Meta decide that actually it wasn't a good algorithmic change at scale? I don't know. It was really interesting to me as somebody who's had people on the podcast to discuss this. It's interesting from the perspective of what's happening in AI right now, but also from the perspective of the fact that I've been having abstract conversations with people about what an intelligence explosion would look like, or what it would look like for AI to automate AI R&D. Just getting a more tangible sense of what's involved in making this AI progress. One of the questions I was debating with Daniel, or I was asking him, is how many of the improvements require a deep conceptual understanding versus how many are just monkeys trying ideas where you could just run a bunch in parallel. It seems like the MLA thing is motivated by this deep conceptual understanding of, “each attention head only needs to see the subspace that's relevant to its attention pattern.” I feel like that just required a lot of conceptual insight in a way that these models are especially bad at. I don't know how the load balancing thing works, but that just seems like maybe you could try it out and see what happens.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.