Evidence receipt / evaluation
Published · transcript-backedNathan Lambert: evaluation
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Lambert
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…Let’s go. Let’s go into the transformer. The transformer is a thing that is talked about a lot, and we will not cover every detail. Essentially the transformer is built on repeated blocks of this attention mechanism and then a traditional dense fully connected multilayer perception, whatever word you want to use for your normal neural network. And you alternate these blocks. There’s other details and where mixture of experts is applied is at this dense model. The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters. We should also say that a transformer is a giant neural network.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.