High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Sholto Douglas: observation

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“Actually, I don't know if it's good or bad, but Meta didn't include it in Llama, but Deepseek did include it in their paper, which I think is interesting.”

— Sholto Douglas

Source trail

Everything needed to verify it.

Speaker
Sholto Douglas
Attribution
Verified speaker
Claim type
observation
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…Does all training involve continuously fucking with these values as you're going through it? It depends on what your architecture is. But I thought it was just cute that you can see them running up into this very hardware-level constraint, trying to go, "What do we wish we could express algorithmically? What can we express under our constraints?" and iteratively solving to get better constraints. And doing this in a really simple and elegant way, and then backing it up with great engineering. I also thought it was interesting that they incorporated the multi-token prediction thing from Meta. So Meta had a nice paper on this multi-token prediction thing. Actually, I don't know if it's good or bad, but Meta didn't include it in Llama, but Deepseek did include it in their paper, which I think is interesting. Was that because they were faster at iterating and including an algorithm? Or did Meta decide that actually it wasn't a good algorithmic change at scale? I don't know. It was really interesting to me as somebody who's had people on the podcast to discuss this. It's interesting from the perspective of what's happening in AI right now, but also from the perspective of the fact that I've been having abstract conversations with people about what an intelligence explosion would look like, or what it would look like for AI to automate AI R&D. Just getting a more tangible sense of what's involved in making this AI progress. One of the questions I was debating with Daniel, or I was asking him, is how many of the improvements require a deep conceptual understanding versus how many are just monkeys trying ideas where you could just run a bunch in parallel. It seems like the MLA thing is motivated by this deep conceptual understanding of, “each attention head only needs to see the subspace that's relevant to its attention pattern.” I feel like that just required a lot of conceptual insight in a way that these models are especially bad at. I don't know how the load balancing thing works, but that just seems like maybe you could try it out and see what happens.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence