High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Dylan Patel

Published podcast speaker

Claims
90
Episodes
7
Shows
3
Named items
0

Claim ledger

What Dylan said.

25 transcript-backed records

01 / evaluation

They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference.

“They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

02 / evaluation

Because they know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.

“Because they know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

03 / evaluation

Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff. So ultimately you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infra side.

“Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff. So ultimately you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infra side.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

04 / evaluation

When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic even though there’s a revenue share and credit back all that.

“When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic even though there’s a revenue share and credit back all that.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

05 / evaluation

We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.

“We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

06 / evaluation

TSMC’s margins on high-performance computing—HPC, AI chips, et cetera—are higher than they are for mobile, because they have a bigger advantage in HPC than they do in mobile.

“TSMC’s margins on high-performance computing—HPC, AI chips, et cetera—are higher than they are for mobile, because they have a bigger advantage in HPC than they do in mobile.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

07 / evaluation

4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.

“4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

08 / evaluation

If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.

“If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

09 / evaluation

You can look across the space at hedge funds and look at their 13Fs and see they own, maybe not exactly what Leopold does, because it’s always a question of what is the most constrained thing.

“You can look across the space at hedge funds and look at their 13Fs and see they own, maybe not exactly what Leopold does, because it’s always a question of what is the most constrained thing.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

10 / evaluation

If you look at PJM, which I think is the largest grid in America—covering the Midwest and some of the Northeast area—in their models they want to have roughly 20 percent excess capacity.

“If you look at PJM, which I think is the largest grid in America—covering the Midwest and some of the Northeast area—in their models they want to have roughly 20 percent excess capacity.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

12 / evaluation

There’s a lot of R&D, there’s a lot of customer acquisition costs. This is sort of why, not Microsoft, but the SaaS companies have underperformed massively in the markets, because the COGS of AI is just so high, and that just completely breaks how these business models work.

“There’s a lot of R&D, there’s a lot of customer acquisition costs. This is sort of why, not Microsoft, but the SaaS companies have underperformed massively in the markets, because the COGS of AI is just so high, and that just completely breaks how these business models work.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

13 / evaluation

You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.

“You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

14 / evaluation

What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.

“What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

15 / evaluation

Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.

“Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.”
Speaker
Dylan Patel
Publisher
Lex Fridman Podcast

18 / evaluation

All in all, you only save about 20% in power per transistor. But because of data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node than just the individual transistors' power efficiency benefit.

“All in all, you only save about 20% in power per transistor. But because of data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node than just the individual transistors' power efficiency benefit.”
Speaker
Dylan Patel
Publisher
Dwarkesh Podcast

20 / evaluation

can ballpark things like, you know, like, yeah, so, so, so, but basically like the, the point is like you're trading this is optimal, this is theoretically the most optimal architecture for performance power and area in a given, and you know, not, not specifically Grok, but VLIW in general is gonna get you closer to optimal there, but then you're giving off, you know, that, that last P, which is pain in the ass program, is, is I think the most simple way to get into it.

“can ballpark things like, you know, like, yeah, so, so, so, but basically like the, the point is like you're trading this is optimal, this is theoretically the most optimal architecture for performance power and area in a given, and you know, not, not specifically Grok, but VLIW in general is gonna get you closer to optimal there, but then you're giving off, you know, that, that last P, which is pain in the ass program, is, is I think the most simple way to get into it.”
Speaker
Dylan Patel
Publisher
Latent Space

21 / evaluation

Building a bigger and better model every, you know, every few months. And I don't know how Apple gets on that train, but you know, at the same time, there's no company that has more powerful distribution, right?

“Building a bigger and better model every, you know, every few months. And I don't know how Apple gets on that train, but you know, at the same time, there's no company that has more powerful distribution, right?”
Speaker
Dylan Patel
Publisher
Latent Space

23 / evaluation

Their chips are not as good, or if they are, even though, you know, I mentioned Intel and AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?

“Their chips are not as good, or if they are, even though, you know, I mentioned Intel and AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?”
Speaker
Dylan Patel
Publisher
Latent Space

24 / evaluation

I mean, it was good. They were doing good, and NVIDIA bought them, you know, in 19, I believe, or 18, but Broadcom has been number one in networking for a decade plus, and Google partnered with them on making the TPU, right?

“I mean, it was good. They were doing good, and NVIDIA bought them, you know, in 19, I believe, or 18, but Broadcom has been number one in networking for a decade plus, and Google partnered with them on making the TPU, right?”
Speaker
Dylan Patel
Publisher
Latent Space

25 / evaluation

I think that one's really critical because it explains Google's infrastructure quite a bit from networking through chips, through all that sort of history of the TPU a little bit.

“I think that one's really critical because it explains Google's infrastructure quite a bit from networking through chips, through all that sort of history of the TPU a little bit.”
Speaker
Dylan Patel
Publisher
Latent Space
Search evidence