High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Lex Fridman Podcast / episode intelligence

#459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

3 Feb 2025 70 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 45evaluation 11uncertainty 8commitment 2prediction 2observation 1recommendation 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

70 published records

01 / evaluation

You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.

“You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

02 / evaluation

The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.

“The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

05 / belief

I think we need to make a point clear on why the time is now for people that don’t think about this, because essentially, with export controls, you’re making it so China cannot make or get cutting edge chips.

“I think we need to make a point clear on why the time is now for people that don’t think about this, because essentially, with export controls, you’re making it so China cannot make or get cutting edge chips.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

06 / belief

I think if you have the demand and the money is on the line, the American companies figure it out. It’s going to take handholding with the government, but I think that the culture helps TSMC break through and it’s easier for them.

“I think if you have the demand and the money is on the line, the American companies figure it out. It’s going to take handholding with the government, but I think that the culture helps TSMC break through and it’s easier for them.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

07 / belief

All of these things can go happen faster. And so I think software and then the other domain is industrial, chemical, mechanical engineers suck at coding just generally.

“All of these things can go happen faster. And so I think software and then the other domain is industrial, chemical, mechanical engineers suck at coding just generally.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

08 / uncertainty

The thing that amplifies the relevance of culture with language models is that we are used to this mode of interacting with people in back and forth conversation. And we now have very powerful computer system that slots into a social context that we’re used to, which makes people very… We don’t know the extent that which people can be impacted by that.

“The thing that amplifies the relevance of culture with language models is that we are used to this mode of interacting with people in back and forth conversation. And we now have very powerful computer system that slots into a social context that we’re used to, which makes people very… We don’t know the extent that which people can be impacted by that.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

09 / belief

I think four of Amazon’s top five revenue products, margin products like gross profit products are all database-related products like Redshift and all these things.

“I think four of Amazon’s top five revenue products, margin products like gross profit products are all database-related products like Redshift and all these things.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

10 / belief

I think also the nature of what it means to be a programmer and what kind of jobs programmers do changes, because I think there needs to be a human in the loop of everything you’ve talked about.

“I think also the nature of what it means to be a programmer and what kind of jobs programmers do changes, because I think there needs to be a human in the loop of everything you’ve talked about.”
Speaker
Lex Fridman
Publisher
Lex Fridman Podcast

11 / belief

There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.

“There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

14 / belief

I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.

“I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

17 / belief

I think to some extent, we have capabilities that hit a certain point where any one person could say, “Oh, okay, if I can leverage those capabilities for X amount of time, this is AGI, call it ’27, ’28.

“I think to some extent, we have capabilities that hit a certain point where any one person could say, “Oh, okay, if I can leverage those capabilities for X amount of time, this is AGI, call it ’27, ’28.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

20 / evaluation

What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.

“What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

21 / belief

I think there’s been some people that are higher level economics understanding say that as you go from 1 billion of smuggling to 10 billion, it’s like you’re hiding certain levels of economic activity and that’s the most reasonable thing to me is that there’s going to be some level where it’s so obvious that it’s easier to find this economic activity.

“I think there’s been some people that are higher level economics understanding say that as you go from 1 billion of smuggling to 10 billion, it’s like you’re hiding certain levels of economic activity and that’s the most reasonable thing to me is that there’s going to be some level where it’s so obvious that it’s easier to find this economic activity.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

22 / belief

There are other international things that are worrying, but there’s just fundamental human goodness and trying to amplify that. I think we’re on a tenuous time.

“There are other international things that are worrying, but there’s just fundamental human goodness and trying to amplify that. I think we’re on a tenuous time.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

23 / uncertainty

” The actual thing that happened is much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade, I don’t know anything about financial markets, but it builds up over the weekend, the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming, but it comes on the weekend and then everybody wants to sell, and then that is a social contagion.

“” The actual thing that happened is much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade, I don’t know anything about financial markets, but it builds up over the weekend, the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming, but it comes on the weekend and then everybody wants to sell, and then that is a social contagion.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

24 / uncertainty

I would say the simplest one is that our language models to date have been designed to give the right answer the highest percentage of the time in one response. And we are now opening the door to different ways of running inference on our models in which we need to reevaluate many parts of the training process, which normally opens the door to more progress, but we don’t know if OpenAI changed a lot or if just sampling more and multiple choice is what they’re doing or if it’s something more complex, but they changed the training and they know that the inference mode is going to be different.

“I would say the simplest one is that our language models to date have been designed to give the right answer the highest percentage of the time in one response. And we are now opening the door to different ways of running inference on our models in which we need to reevaluate many parts of the training process, which normally opens the door to more progress, but we don’t know if OpenAI changed a lot or if just sampling more and multiple choice is what they’re doing or if it’s something more complex, but they changed the training and they know that the inference mode is going to be different.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

25 / evaluation

Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.

“Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

29 / uncertainty

I think OpenAI’s statement, I don’t know if you’ve seen the five levels where it’s chat is level one, reasoning is level two, and then agents is level three.

“I think OpenAI’s statement, I don’t know if you’ve seen the five levels where it’s chat is level one, reasoning is level two, and then agents is level three.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

30 / observation

Character AI very likely could be optimizing this where it’s the way that this data is collected is naive, whereas you’re presented a few options and you choose them. But that’s not the only way that these models are going to be trained.

“Character AI very likely could be optimizing this where it’s the way that this data is collected is naive, whereas you’re presented a few options and you choose them. But that’s not the only way that these models are going to be trained.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

32 / belief

China will win because of these restrictions long-term, unless AI does something in the short-term, which I believe AI will make massive changes to society in the medium, short-term.

“China will win because of these restrictions long-term, unless AI does something in the short-term, which I believe AI will make massive changes to society in the medium, short-term.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

38 / belief

I think if you look at every layer of the compute stack, whether it goes from lithography and etch all the way to fabrication, to optics, to networking, to power, to transformers, to cooling, to a networking, and you just go on up and up and up and up the stack, even air conditioners for data centers are innovating.

“I think if you look at every layer of the compute stack, whether it goes from lithography and etch all the way to fabrication, to optics, to networking, to power, to transformers, to cooling, to a networking, and you just go on up and up and up and up the stack, even air conditioners for data centers are innovating.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

39 / belief

I think, and there were a lot of false narratives, which is like, “Hey, these guys are spending billions on models,” and they’re not spending billions on models.

“I think, and there were a lot of false narratives, which is like, “Hey, these guys are spending billions on models,” and they’re not spending billions on models.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

43 / belief

Of what GenAI can do to society, but it was very clear, I think, to at least National Security Council and those sort of folks, that this was where the world is headed, this cold war that’s happening.

“Of what GenAI can do to society, but it was very clear, I think, to at least National Security Council and those sort of folks, that this was where the world is headed, this cold war that’s happening.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

45 / belief

I think the clearest example we have, because Meta is also open, they talk about order of 60k to 100k H100 equivalent GPUs in their training clusters.

“I think the clearest example we have, because Meta is also open, they talk about order of 60k to 100k H100 equivalent GPUs in their training clusters.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

49 / belief

I think we already have something like that’s exceptionally impressive that Alan Turing would for sure say is AGI, but he’s referring more to something once in possession of, then you would have a significant military and geopolitical advantage over other nations.

“I think we already have something like that’s exceptionally impressive that Alan Turing would for sure say is AGI, but he’s referring more to something once in possession of, then you would have a significant military and geopolitical advantage over other nations.”
Speaker
Lex Fridman
Publisher
Lex Fridman Podcast

50 / belief

I would say that the long tail of use is going to go inside of AI, which is if you scrape trillions of tokens of data, you’re not looking and saying, “This one New York Times article is so important to me.

“I would say that the long tail of use is going to go inside of AI, which is if you scrape trillions of tokens of data, you’re not looking and saying, “This one New York Times article is so important to me.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

51 / uncertainty

I think for reasoning with this RL and verifiable domains, we’re early, but we don’t know where the point is where you just start training on enough domains and poof, more domains just start working.

“I think for reasoning with this RL and verifiable domains, we’re early, but we don’t know where the point is where you just start training on enough domains and poof, more domains just start working.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

52 / belief

We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.

“We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

54 / belief

For the past few years, the highest cost human data has been in these preferences, which is comparing, I would say, highest cost and highest total usage, so a lot of money has gone to these pairwise comparisons where you have two model outputs and a human is comparing between the two of them.

“For the past few years, the highest cost human data has been in these preferences, which is comparing, I would say, highest cost and highest total usage, so a lot of money has gone to these pairwise comparisons where you have two model outputs and a human is comparing between the two of them.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

57 / belief

I think when you look at the Chinese labs, Huawei has a lab, Moonshot AI, there’s a couple other labs out there that are really close with the government, and then there’s labs like Alibaba and DeepSeek, which are not close with the government.

“I think when you look at the Chinese labs, Huawei has a lab, Moonshot AI, there’s a couple other labs out there that are really close with the government, and then there’s labs like Alibaba and DeepSeek, which are not close with the government.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

58 / belief

Generally, humans have positive impacts on the world, at least societally, but it’s possible for individual humans to have such negative impacts. And AGI, at least as I think the labs define it, which is not a runaway sentient thing, but rather just something that can do a lot of tasks really efficiently amplifies the capabilities of someone causing extreme damage.

“Generally, humans have positive impacts on the world, at least societally, but it’s possible for individual humans to have such negative impacts. And AGI, at least as I think the labs define it, which is not a runaway sentient thing, but rather just something that can do a lot of tasks really efficiently amplifies the capabilities of someone causing extreme damage.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

59 / belief

I think in terms of internet posts and things that people have been measuring, it hasn’t been a exponential increase or something extremely measurable and things you’re talking about with voice calls and stuff like that, it could be in modalities that are harder to measure.

“I think in terms of internet posts and things that people have been measuring, it hasn’t been a exponential increase or something extremely measurable and things you’re talking about with voice calls and stuff like that, it could be in modalities that are harder to measure.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

60 / belief

Very long-term motivated in how the ecosystem of AI should work. And I think from a Chinese perspective, he wants a Chinese company to build this vision.

“Very long-term motivated in how the ecosystem of AI should work. And I think from a Chinese perspective, he wants a Chinese company to build this vision.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

61 / evaluation

Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.

“Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

63 / evaluation

It’s insane. It’s exciting because they’re trying to figure out what the bottlenecks are, how to remove the bottlenecks, how to make sure that… There’s just so many really cool things about putting together a data center because everything has to work.

“It’s insane. It’s exciting because they’re trying to figure out what the bottlenecks are, how to remove the bottlenecks, how to make sure that… There’s just so many really cool things about putting together a data center because everything has to work.”
Speaker
Lex Fridman
Publisher
Lex Fridman Podcast

64 / evaluation

If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.

“If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

65 / commitment

We want to build the AI system that people around the world are using. And I think that, if anything, some of the recent news has only strengthened our conviction that this is the right thing to be focused on.

“We want to build the AI system that people around the world are using. And I think that, if anything, some of the recent news has only strengthened our conviction that this is the right thing to be focused on.”
Speaker
Lex Fridman
Publisher
Lex Fridman Podcast

66 / prediction

I think we should summarize what The Bitter Lesson actually is about, is that The Bitter Lesson essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that which are scalable in learning and search, is what it calls out.

“I think we should summarize what The Bitter Lesson actually is about, is that The Bitter Lesson essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that which are scalable in learning and search, is what it calls out.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

67 / prediction

I think mostly if you’re going to say that I’m feeling the AGI is that I expect continued, rapid, surprising progress over the next few years. So, something like R1 is less surprising to me from DeepSeek because I expect there to be new paradigms versus … … surprising to me from DeepSeek because I expect there to be new paradigms where substantial progress can be made.

“I think mostly if you’re going to say that I’m feeling the AGI is that I expect continued, rapid, surprising progress over the next few years. So, something like R1 is less surprising to me from DeepSeek because I expect there to be new paradigms versus … … surprising to me from DeepSeek because I expect there to be new paradigms where substantial progress can be made.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

68 / evaluation

There’s not full agreement in the community, but for us that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.

“There’s not full agreement in the community, but for us that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

69 / evaluation

Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.

“Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

70 / evaluation

We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.

“We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast
Search evidence