High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Dylan Patel

Published podcast speaker

Claims
90
Episodes
7
Shows
3
Named items
0

Claim ledger

What Dylan said.

90 transcript-backed records

02 / prediction

Next year, a big new entrant is, for example, SpaceX, which is building a ton of compute. They’re actively going to lease quite a bit of it to Anthropic and OpenAI, most likely, because they’re the ones who have the marginal capability to pay the highest price.

“Next year, a big new entrant is, for example, SpaceX, which is building a ton of compute. They’re actively going to lease quite a bit of it to Anthropic and OpenAI, most likely, because they’re the ones who have the marginal capability to pay the highest price.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

05 / evaluation

They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference.

“They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

06 / evaluation

Because they know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.

“Because they know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

07 / observation

You’ve seen people do funny arbitrages here where they buy turbines and then try and resell them, because the value of a turbine is way more since it’s the thing bottlenecking your data center.

“You’ve seen people do funny arbitrages here where they buy turbines and then try and resell them, because the value of a turbine is way more since it’s the thing bottlenecking your data center.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

08 / prediction

Anthropic not releasing what their safety assessment says is Model 2, which is widely believed to be the next version of Mythos. They’re clearly not releasing their best models, in which case their revenue per megawatt stalls or can even start to decline again because other models are competitive again.

“Anthropic not releasing what their safety assessment says is Model 2, which is widely believed to be the next version of Mythos. They’re clearly not releasing their best models, in which case their revenue per megawatt stalls or can even start to decline again because other models are competitive again.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

09 / evaluation

Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff. So ultimately you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infra side.

“Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff. So ultimately you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infra side.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

11 / evaluation

When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic even though there’s a revenue share and credit back all that.

“When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic even though there’s a revenue share and credit back all that.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

12 / evaluation

We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.

“We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

13 / evaluation

TSMC’s margins on high-performance computing—HPC, AI chips, et cetera—are higher than they are for mobile, because they have a bigger advantage in HPC than they do in mobile.

“TSMC’s margins on high-performance computing—HPC, AI chips, et cetera—are higher than they are for mobile, because they have a bigger advantage in HPC than they do in mobile.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

14 / evaluation

4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.

“4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

15 / evaluation

If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.

“If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

18 / observation

The problem is that getting the heat out of that dense area means you have to move away from standard air and liquid cooling to more exotic forms of liquid cooling, or even immersion, to get to higher power densities.

“The problem is that getting the heat out of that dense area means you have to move away from standard air and liquid cooling to more exotic forms of liquid cooling, or even immersion, to get to higher power densities.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

19 / belief

I think going back to the earlier view that if the models are so powerful, the value of a GPU goes up over time, right now only OpenAI and Anthropic have that viewpoint.

“I think going back to the earlier view that if the models are so powerful, the value of a GPU goes up over time, right now only OpenAI and Anthropic have that viewpoint.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

22 / uncertainty

Then you’ll have this insane margin that ASML and TSMC should have been charging. But the thing is, I don’t know if ASML and TSMC will ever agree to this.

“Then you’ll have this insane margin that ASML and TSMC should have been charging. But the thing is, I don’t know if ASML and TSMC will ever agree to this.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

24 / prediction

In some cases, the argument people are making is if you didn’t sign a long-term deal, because every two years NVIDIA is tripling or quadrupling the performance while only 2X-ing or 50% increasing the price… Then the price of an H100… Sure maybe the value in the market was $2 at 35% gross margins in 2024, but in 2026, when Blackwell is in super high volume and deploying millions a year, you’re actually now worth $1/hour.

“In some cases, the argument people are making is if you didn’t sign a long-term deal, because every two years NVIDIA is tripling or quadrupling the performance while only 2X-ing or 50% increasing the price… Then the price of an H100… Sure maybe the value in the market was $2 at 35% gross margins in 2024, but in 2026, when Blackwell is in super high volume and deploying millions a year, you’re actually now worth $1/hour.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

25 / evaluation

You can look across the space at hedge funds and look at their 13Fs and see they own, maybe not exactly what Leopold does, because it’s always a question of what is the most constrained thing.

“You can look across the space at hedge funds and look at their 13Fs and see they own, maybe not exactly what Leopold does, because it’s always a question of what is the most constrained thing.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

28 / prediction

Even if you halve smartphone volumes, because of the shape of the halving, the low end gets cut by more than half, while the high end gets cut by less than half, because you and I will still buy the high-end phones that cost north of a thousand dollars.

“Even if you halve smartphone volumes, because of the shape of the halving, the low end gets cut by more than half, while the high end gets cut by less than half, because you and I will still buy the high-end phones that cost north of a thousand dollars.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

29 / evaluation

If you look at PJM, which I think is the largest grid in America—covering the Midwest and some of the Northeast area—in their models they want to have roughly 20 percent excess capacity.

“If you look at PJM, which I think is the largest grid in America—covering the Midwest and some of the Northeast area—in their models they want to have roughly 20 percent excess capacity.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

31 / evaluation

There’s a lot of R&D, there’s a lot of customer acquisition costs. This is sort of why, not Microsoft, but the SaaS companies have underperformed massively in the markets, because the COGS of AI is just so high, and that just completely breaks how these business models work.

“There’s a lot of R&D, there’s a lot of customer acquisition costs. This is sort of why, not Microsoft, but the SaaS companies have underperformed massively in the markets, because the COGS of AI is just so high, and that just completely breaks how these business models work.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

32 / evaluation

You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.

“You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

34 / belief

All of these things can go happen faster. And so I think software and then the other domain is industrial, chemical, mechanical engineers suck at coding just generally.

“All of these things can go happen faster. And so I think software and then the other domain is industrial, chemical, mechanical engineers suck at coding just generally.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

35 / belief

I think four of Amazon’s top five revenue products, margin products like gross profit products are all database-related products like Redshift and all these things.

“I think four of Amazon’s top five revenue products, margin products like gross profit products are all database-related products like Redshift and all these things.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

38 / belief

I think to some extent, we have capabilities that hit a certain point where any one person could say, “Oh, okay, if I can leverage those capabilities for X amount of time, this is AGI, call it ’27, ’28.

“I think to some extent, we have capabilities that hit a certain point where any one person could say, “Oh, okay, if I can leverage those capabilities for X amount of time, this is AGI, call it ’27, ’28.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

40 / evaluation

What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.

“What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

41 / evaluation

Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.

“Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

43 / uncertainty

I think OpenAI’s statement, I don’t know if you’ve seen the five levels where it’s chat is level one, reasoning is level two, and then agents is level three.

“I think OpenAI’s statement, I don’t know if you’ve seen the five levels where it’s chat is level one, reasoning is level two, and then agents is level three.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

44 / belief

China will win because of these restrictions long-term, unless AI does something in the short-term, which I believe AI will make massive changes to society in the medium, short-term.

“China will win because of these restrictions long-term, unless AI does something in the short-term, which I believe AI will make massive changes to society in the medium, short-term.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

48 / belief

I think if you look at every layer of the compute stack, whether it goes from lithography and etch all the way to fabrication, to optics, to networking, to power, to transformers, to cooling, to a networking, and you just go on up and up and up and up the stack, even air conditioners for data centers are innovating.

“I think if you look at every layer of the compute stack, whether it goes from lithography and etch all the way to fabrication, to optics, to networking, to power, to transformers, to cooling, to a networking, and you just go on up and up and up and up the stack, even air conditioners for data centers are innovating.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

49 / belief

I think, and there were a lot of false narratives, which is like, “Hey, these guys are spending billions on models,” and they’re not spending billions on models.

“I think, and there were a lot of false narratives, which is like, “Hey, these guys are spending billions on models,” and they’re not spending billions on models.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

52 / belief

Of what GenAI can do to society, but it was very clear, I think, to at least National Security Council and those sort of folks, that this was where the world is headed, this cold war that’s happening.

“Of what GenAI can do to society, but it was very clear, I think, to at least National Security Council and those sort of folks, that this was where the world is headed, this cold war that’s happening.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

55 / belief

I think when you look at the Chinese labs, Huawei has a lab, Moonshot AI, there’s a couple other labs out there that are really close with the government, and then there’s labs like Alibaba and DeepSeek, which are not close with the government.

“I think when you look at the Chinese labs, Huawei has a lab, Moonshot AI, there’s a couple other labs out there that are really close with the government, and then there’s labs like Alibaba and DeepSeek, which are not close with the government.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

56 / belief

Generally, humans have positive impacts on the world, at least societally, but it’s possible for individual humans to have such negative impacts. And AGI, at least as I think the labs define it, which is not a runaway sentient thing, but rather just something that can do a lot of tasks really efficiently amplifies the capabilities of someone causing extreme damage.

“Generally, humans have positive impacts on the world, at least societally, but it’s possible for individual humans to have such negative impacts. And AGI, at least as I think the labs define it, which is not a runaway sentient thing, but rather just something that can do a lot of tasks really efficiently amplifies the capabilities of someone causing extreme damage.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

62 / belief

I think more reasonably, it will be like the total summation of the flops that you deliver to the model across pre-training, post-training, synthetic data for that pre-training and post-training data, as well as some of the inference time compute efficiencies.

“I think more reasonably, it will be like the total summation of the flops that you deliver to the model across pre-training, post-training, synthetic data for that pre-training and post-training data, as well as some of the inference time compute efficiencies.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

68 / evaluation

All in all, you only save about 20% in power per transistor. But because of data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node than just the individual transistors' power efficiency benefit.

“All in all, you only save about 20% in power per transistor. But because of data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node than just the individual transistors' power efficiency benefit.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

73 / recommendation

If you want to restrict them from having chips, you have to let them have at least some level of chip that is better than what they can build internally.

“If you want to restrict them from having chips, you have to let them have at least some level of chip that is better than what they can build internally.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

78 / uncertainty

I mean, I don't know if I, if I have a coherent thesis, but it's, it's sure fun to, it's Who, who think that like, I, I, I just have an intense hatred for RAG.

“I mean, I don't know if I, if I have a coherent thesis, but it's, it's sure fun to, it's Who, who think that like, I, I, I just have an intense hatred for RAG.”
Speaker
Dylan Patel
Publisher
Latent Space

79 / evaluation

can ballpark things like, you know, like, yeah, so, so, so, but basically like the, the point is like you're trading this is optimal, this is theoretically the most optimal architecture for performance power and area in a given, and you know, not, not specifically Grok, but VLIW in general is gonna get you closer to optimal there, but then you're giving off, you know, that, that last P, which is pain in the ass program, is, is I think the most simple way to get into it.

“can ballpark things like, you know, like, yeah, so, so, so, but basically like the, the point is like you're trading this is optimal, this is theoretically the most optimal architecture for performance power and area in a given, and you know, not, not specifically Grok, but VLIW in general is gonna get you closer to optimal there, but then you're giving off, you know, that, that last P, which is pain in the ass program, is, is I think the most simple way to get into it.”
Speaker
Dylan Patel
Publisher
Latent Space

80 / belief

Like, you look at what Tsinghua University is doing in China, actually, they open sourced their model to I think the largest like by parameter count, at least open source models.

“Like, you look at what Tsinghua University is doing in China, actually, they open sourced their model to I think the largest like by parameter count, at least open source models.”
Speaker
Dylan Patel
Publisher
Latent Space

81 / belief

I think with AI, especially like how scaling laws are going, it's like incredibly important for infrastructure is like so much more important. And then like when you just think about software cost, right, like the cost structure of it, there was always a bigger component of R&D and like SAS businesses, you know, all over SF, all these SAS businesses did crazy good because, you know, they just start as they grow and then all of a sudden they're so freaking profitable for each incremental new customer.

“I think with AI, especially like how scaling laws are going, it's like incredibly important for infrastructure is like so much more important. And then like when you just think about software cost, right, like the cost structure of it, there was always a bigger component of R&D and like SAS businesses, you know, all over SF, all these SAS businesses did crazy good because, you know, they just start as they grow and then all of a sudden they're so freaking profitable for each incremental new customer.”
Speaker
Dylan Patel
Publisher
Latent Space

82 / evaluation

Building a bigger and better model every, you know, every few months. And I don't know how Apple gets on that train, but you know, at the same time, there's no company that has more powerful distribution, right?

“Building a bigger and better model every, you know, every few months. And I don't know how Apple gets on that train, but you know, at the same time, there's no company that has more powerful distribution, right?”
Speaker
Dylan Patel
Publisher
Latent Space

85 / evaluation

Their chips are not as good, or if they are, even though, you know, I mentioned Intel and AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?

“Their chips are not as good, or if they are, even though, you know, I mentioned Intel and AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?”
Speaker
Dylan Patel
Publisher
Latent Space

86 / evaluation

I mean, it was good. They were doing good, and NVIDIA bought them, you know, in 19, I believe, or 18, but Broadcom has been number one in networking for a decade plus, and Google partnered with them on making the TPU, right?

“I mean, it was good. They were doing good, and NVIDIA bought them, you know, in 19, I believe, or 18, but Broadcom has been number one in networking for a decade plus, and Google partnered with them on making the TPU, right?”
Speaker
Dylan Patel
Publisher
Latent Space

87 / recommendation

Unless you're fine tuning for on-device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time because the cost of inference is actually much cheaper than you think once you achieve good MBU and you batch at a decent size, which any successful business in the cloud is going to achieve, you know, and then two, fine tuning like people like, oh, you know, this 7 billion parameter model, if you fine tune it on a data set is almost as good as 3.

“Unless you're fine tuning for on-device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time because the cost of inference is actually much cheaper than you think once you achieve good MBU and you batch at a decent size, which any successful business in the cloud is going to achieve, you know, and then two, fine tuning like people like, oh, you know, this 7 billion parameter model, if you fine tune it on a data set is almost as good as 3.”
Speaker
Dylan Patel
Publisher
Latent Space

88 / preference

Some people would argue lower, right? But at the very least, you need to achieve human reading level speeds and probably a little bit faster, because we like their skin, to have a usable model for chatbot style applications.

“Some people would argue lower, right? But at the very least, you need to achieve human reading level speeds and probably a little bit faster, because we like their skin, to have a usable model for chatbot style applications.”
Speaker
Dylan Patel
Publisher
Latent Space

89 / evaluation

I think that one's really critical because it explains Google's infrastructure quite a bit from networking through chips, through all that sort of history of the TPU a little bit.

“I think that one's really critical because it explains Google's infrastructure quite a bit from networking through chips, through all that sort of history of the TPU a little bit.”
Speaker
Dylan Patel
Publisher
Latent Space

90 / prediction

So in training, everyone just talks about MFU, right? But then on inference, right, which I think is one LLM inference will be bigger than training or multimodal, whatever, bubble inference will be bigger than training, probably next year, in fact, at least in terms of GPUs deployed.

“So in training, everyone just talks about MFU, right? But then on inference, right, which I think is one LLM inference will be bigger than training or multimodal, whatever, bubble inference will be bigger than training, probably next year, in fact, at least in terms of GPUs deployed.”
Speaker
Dylan Patel
Publisher
Latent Space
Search evidence