Evidence receipt / belief
Published · transcript-backedDylan Patel: belief
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“I think there’s a couple factors here. One is that they do have model architecture innovations.”
Source trail
Everything needed to verify it.
- Speaker
- Dylan Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…Yeah, let’s talk about why it’s so cheap on the inference. It works well and it’s cheap. Why is R1 so damn cheap? I think there’s a couple factors here. One is that they do have model architecture innovations. This MLA, this new attention that they’ve done, is different than the attention from attention is all you need, the transformer attention. Now, others have already innovated. There’s a lot of work like MQA, GQA, local, global, all these different innovations that try to bend the curve. It’s still quadratic, but the constant is now smaller. Related to our previous discussion, this multi-head latent attention can save about 80 to 90% in memory from the attention mechanism, which helps especially in long contexts.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.