Evidence receipt / belief
Published · transcript-backedDylan Patel: belief
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“Better lesson, right? The question is, I think, when not if, because the rate of progress is so fast.”
Source trail
Everything needed to verify it.
- Speaker
- Dylan Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…So you have just a linear decline on log scale from GPT-3 through 3.5 to Lama- It’s like five cents or something like that now, right? Versus $60, 1200X, that’s not the exact numbers, but it’s 1200X, I remember that number, is humongous cost per intelligence. Now, the freak out over DeepSeek is, “Oh my god, they made it so cheap.” It’s like actually, if you look at this trend line, they’re not below the trend line first of all, at least for GPT-3, right? They are the first to hit it, which is a big deal, but they’re not below the trend line as far as GPT-3. Now we have GPT-4, what’s going to happen with these reasoning capabilities? It’s a mix of architectural innovations, it’s a mix of better data, and it’s going to be better training techniques and all of these better inference systems, better hardware going from each generation of GPU to new generations or ASICs. Everything is going to take this cost curve down and down and down and down. And then can I just spawn a thousand different LLMs to create a task and then pick from one of them? Or whatever search technique, I want, a Tree, Monte Carlo Tree Search, maybe it gets that complicated, maybe it doesn’t because it’s too complicated to actually scale. Who knows? Better lesson, right? The question is, I think, when not if, because the rate of progress is so fast. Nine months ago, Dario said nine months ago the cost to train an inference was this, and now we’re much better than this and DeepSeek is much better than this. And that cost curve for GPT-4, which was also roughly $60 per million tokens when it launched, has already fallen to $2 or so. And we’re going to get it down to cents probably for GPT-4 quality. And then that’s the base for the reasoning models like o1 that we have today and o1-pro is spawning multiple and o3 and so on and so forth, these search techniques, too expensive today, but they will get cheaper and that’s what’s going to unlock the intelligence. So, it’ll get cheaper and cheaper and cheaper. The big DeepSeek R1 release freaked everybody out because of the cheaper. One of the manifestations of that is NVIDIA stock plummeted. Can you explain what happened? And also just explain this moment and if NVIDIA is going to keep winning.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.