Evidence receipt / evaluation
Published · transcript-backedDylan Patel: evaluation
13 Mar 2026 Dwarkesh Podcast Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute
“If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.”
Source trail
Everything needed to verify it.
- Speaker
- Dylan Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Mar 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…No, because they both increased by $1. The whole effect is that if there’s a fixed cost that is applied to both. Then the price difference between them, the ratio, changes. Previously, the more expensive one was 2X more expensive. Now it’s just 1.5X more expensive. So I wonder if applied to AI that would mean that, if GPUs are going to get more expensive, there will be a fixed cost increase in the price of compute. As a result, that will push people to be willing to pay higher margins for slightly better models. Because the calculus is, I’m going to be paying all this money for the compute anyway. I might as well just pay slightly more to make sure it’s the very best model rather than a model that’s slightly worse. So the Hopper went from $2 to $3. If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3. Interesting. I think that makes a ton of sense. We just see all of the volumes are on the best models today, all the revenue is on the best models today. In a compute-limited world, two things happen. One, companies that don’t have commitment issues and have these five-year contracts for compute have locked in a humongous margin advantage. They’ve locked in compute for five years at the price it transacted at two, three, or five years ago. Whereas if you’re three years into that five-year contract and someone else’s two-year or three-year contract rolled off, and now they’re trying to buy that at modern pricing, when it’s priced to the value of models, the price is going to be up a lot more. So the person who committed early has better margins in general. The percentage of the market that is in long-term contracts is much larger than the percentage of the market in short-term contracts that can be this flex capacity you add at the last second. At the same time, where does the margin go? Because models get more valuable, how much can the cloud players flex their pricing? If you look at CoreWeave, their average term duration is over three years right now. For ninety-eight percent plus of their compute, it’s over three years. They end up with this conundrum where they can’t actually flex price. But every year they’re adding incrementally way more capacity than they had previously. This year alone, Meta’s adding as much capacity as they had in their entire fleet of compute and data centers for all purposes for serving WhatsApp, Instagram, and Facebook in 2022, and doing AI. They’re adding that alone this year. In the same sense, you talk about Meta doing that, CoreWeave, Google, and Amazon, all these companies are adding insane amounts of compute year on year. That new compute gets transacted at the new price. In a sense, yes, you’ve locked in, as long as we’re in a takeoff. “Oh, OpenAI went from six hundred megawatts to two gigawatts last year, and from two gigawatts to six plus this year, and six to twelve next year.” The incremental added compute is where all the cost is, not the prior long-term contracts. Then who holds the cards is the infra providers for charging margin. Now the cloud players, the neoclouds, or the hyperscalers can charge the margin. They can to some extent, but then as you go upstream to who has access to all the memory and logic capacity, it’s Nvidia for the most part. They’ve signed a lot of long-term contracts. They’ve got ninety billion dollars of long-term contracts today, and they’re negotiating three-year deals today with the memory vendors. You’ve got Amazon and Google through Broadcom, Amazon directly, and AMD. ot ninety billion dollars of long-term contracts today, and they’re negotiating three-year deals today with the memory vendors. You’ve got Amazon and Google through Broadcom, Amazon directly, and AMD. These companies hold all the cards because they’ve secured the capacity. TSMC is not raising prices, but memory vendors are, to some extent, raising a lot of price. They’re going to double or triple price again, but then they’re also signing these long-term deals. Who is able to accrue all the margin dollars is potentially the cloud, potentially the chip vendors, and the memory vendors, until TSMC or ASML break out and say, “No, we’re going to charge a lot more.” But at the same time, do the model vendors get to charge crazy margins? At least this year, we’re going to see margins for the model vendors go up a lot. Because they’re so capacity constrained, they have to destroy demand. There’s no way Anthropic can continue at the current pace without destroying demand.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.