High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Lambert: evaluation

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.”

— Nathan Lambert

Source trail

Everything needed to verify it.

Speaker
Nathan Lambert
Attribution
Verified speaker
Claim type
evaluation
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…And therefore your serving costs are lower, because the server costs the same. This is eight H100s, roughly $2 an hour per GPU. That’s $16 an hour. That is somewhat of a fixed cost. You can do things to make it lower of course, but it’s like $16 an hour. Now how many users can you serve, how many tokens can you generate, and then you divide the two and that’s your cost. And so with reasoning models, this is where a lot of the complexity comes about and why memory is so important. Because if you have limited amounts of memory, then you can’t serve so many users. If you have limited amounts of memory, your serving speeds get lower. And so your costs get a lot, lot worse because all of a sudden if I was used to, hey, on this $16 an hour server I’m serving Llama 405B, or if I’m serving DeepSeek-V3 and it’s all chat style applications, i.e. we’re just chit-chatting, the sequence length are a thousand, a few thousand. When you use a language model, it’s a few thousand context length most of times. Sometimes you’re dropping a big document, but then you process it, you get your answer, you throw it away, you move on to the next thing. Whereas with reasoning, I’m now generating tens of thousands of tokens in sequence. And so this memory, this KV cache, has to stay resonant and you have to keep loading it, you have to keep it in-memory constantly. And now this butts out other users. If there’s now a reasoning task and the model’s capable of reasoning, then all of a sudden that memory pressure means that I can’t serve as many users simultaneously. Let’s go into DeepSeek again. So we’re in the post DeepSeek-R1 time I think, and there’s two sides to this market, watching how hard it is to serve it. On one side we’re going to talk about DeepSeek themselves. They now have a chat app that got to number one on the App Store. Disclaimer number one on the App Store is measured by velocity, so it’s not necessarily saying that more people have the DeepSeek app than the ChatGPT app. But it is still remarkable. Claude has never hit the number one in the App Store, even though everyone in San Francisco is like, “Oh my god, you got to use Claude. Don’t use ChatGPT.” So DeepSeek hit this. They also launched an API product recently where you can ping their API and get these super long responses for R1 out. At the same time as these are out, we’ll get to what’s happened to them. Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users. We are trying to evaluate R1 because we have really similar research going on. We released the model and we’re trying to compare to it. And out of all the companies that are, quote unquote, “serving” R1 and they’re doing it at prices that are way higher than the DeepSeek API, most of them barely work and the throughput is really low. To give context, one of the parts of freaking us out was like China reached capabilities. The other aspect is they did it so cheap. And the so cheap, we talked about on the training side why it was so cheap slash-…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence