High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Lambert: belief

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“I think a lot of the AI industry is going through this challenge of communications right now where OpenAI makes fun of their own naming schemes.”

— Nathan Lambert

Source trail

Everything needed to verify it.

Speaker
Nathan Lambert
Attribution
Verified speaker
Claim type
belief
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…m many perspectives in this conversation. This is the Lex Fridman podcast, to support it please check out our sponsors in the description. And now, dear friends, here’s Dylan Patel and Nathan Lambert. A lot of people are curious to understand China’s DeepSeek AI models, so let’s lay it out. Nathan, can you describe what DeepSeek-V3 and DeepSeek-R1 are, how they work, how they’re trained? Let’s look at the big picture and then we’ll zoom in on the details. DeepSeek-V3 is a new mixture of experts, transformer language model from DeepSeek who is based in China. They have some new specifics in the model that we’ll get into. Largely this is a open weight model and it’s a instruction model like what you would use in ChatGPT. They also released what is called the base model, which is before these techniques of post-training. Most people use instruction models today, and those are what’s served in all sorts of applications. This was released on, I believe, December 26th or that week. And then weeks later on January 20th, DeepSeek released DeepSeek-R1, which is a reasoning model, which really accelerated a lot of this discussion. This reasoning model has a lot of overlapping training steps to DeepSeek-V3, and it’s confusing that you have a base model called V3 that you do something to to get a chat model and then you do some different things to get a reasoning model. I think a lot of the AI industry is going through this challenge of communications right now where OpenAI makes fun of their own naming schemes. They have GPT-4o, they have OpenIA o1, and there’s a lot of types of models, so we’re going to break down what each of them are. There’s a lot of technical specifics on training and go through them high level to specific and go through each of them. There’s so many places we can go here, but maybe let’s go to open weights first. What does it mean for a model to be open weights and what are the different flavors of open source in general?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence