Evidence receipt / evaluation
Published · transcript-backedDylan Patel: evaluation
13 Mar 2026 Dwarkesh Podcast Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute
“4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.”
Source trail
Everything needed to verify it.
- Speaker
- Dylan Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Mar 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…There’s a few strings to pull on there. One is, what happens to depreciation of GPUs? I guess I didn’t answer your prior question, which is that I think Anthropic will be able to get to five gigawatts-ish, maybe a little bit more by the end of the year through themselves as well as their product being served through Bedrock, Vertex, or Foundry. I think they’ll be able to get to five or six gigawatts, which is way above their initial plans. OpenAI will be roughly the same, actually a little bit higher based on our numbers. But anyway, the depreciation cycle of a GPU. Michael Burry was saying it’s three years or less. That’s sort of his argument. There are two lenses to look at this. Mechanically, there’s a TCO model, total cost of ownership of a GPU, where we project pricing out for GPUs and build up the total cost of a cluster. There are a number of costs: your data center cost, your networking cost, your smart hands and people in the data center swapping stuff out. There’s your spare parts, your actual chip cost, your server cost. All these various costs get lumped together. There’s some depreciation cycles on it, certain credit costs on it. You build up to, “Hey, an H100 costs $1.40/hour to deploy at volume across five years if your depreciation is five years.” If you sign a deal at $2/hour for those five years, your gross margin is roughly 35%. It’s a little bit above that. If you sign it for $1.90, it’s 35% roughly. Then you assume at that fifth year, the GPU falls off a bus and is dead. In some cases, the argument people are making is if you didn’t sign a long-term deal, because every two years NVIDIA is tripling or quadrupling the performance while only 2X-ing or 50% increasing the price… Then the price of an H100… Sure maybe the value in the market was $2 at 35% gross margins in 2024, but in 2026, when Blackwell is in super high volume and deploying millions a year, you’re actually now worth $1/hour. And when Rubin in ‘27 is in super high volume—even though it starts shipping this year, it’s super high volume next year—doing millions of chips a year deployed into clouds, you’ve got another 3X in performance, another 50% or 2X in price, then the Hopper is only worth $0.70/hour. So the price of a GPU would continue to fall. That’s one lens. The other lens is, what is the utility you get out of the chip? If you could build infinite Rubin or infinite of the newest chip, then yes, that’s exactly what would happen. The price of a Hopper would fall at a spot or short-term contract rate as the new chips come out and the price per performance goes up. But because you are so limited on semiconductors and deployment timelines, what actually prices these chips is not the comparative thing I can buy today, but rather what is the value I can derive out of this chip today. In that sense, let’s take GPT-5.4. GPT-5.4 is both way cheaper to run than GPT-4 and has fewer active parameters. ing I can buy today, but rather what is the value I can derive out of this chip today. In that sense, let’s take GPT-5.4. GPT-5. 4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE. There’s also been so many other advancements in training, RL, model architecture, and data qualities that have made GPT-5.4 way better than GPT-4. And it’s cheaper to serve. When you look at an H100, it can serve more tokens per GPU of 5.4 than if you had ran GPT-4 on it. So it’s producing more tokens of a model that is of higher quality. What is the maximum TAM for GPT-4 tokens? Maybe it was a few billion dollars, maybe it was tens of billions of dollars. Adoption takes time. For GPT-5.4, that number is probably north of a hundred billion. But there’s an adoption lag, there’s competition, and there’s the constant improvements that everyone else is having. If improvements stopped here, the value of an H100 is now predicated on the value that GPT-5.4 can get out of it instead of the value that GPT-4 can get out of it. These labs are in a competitive environment, so their margins can’t go to infinity. You sort of have this dynamic that is quite interesting in that an H100 is worth more today than it was three years ago. That’s crazy. It’s also interesting from the perspective of just taking that forward. If we had actual AGI models developed, if we had a genuine human on a server… These are such hand wave-y numbers about how many flops the brain can do. But on a flop basis, an H100 is estimated to do 1e15, which is how much some people estimate the human brain does in flops. Obviously, in terms of memory, the human brain has way more. An H100 is 80 gigabytes, and the brain might have petabytes.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.