01 / prediction
I don’t think it’ll happen the year after, because the world is capital constrained.
“I don’t think it’ll happen the year after, because the world is capital constrained.”
- Speaker
- Dylan Patel
- Publisher
- Dwarkesh Podcast
Public evidence record
Published podcast speaker
Claim ledger
90 transcript-backed records
01 / prediction
“I don’t think it’ll happen the year after, because the world is capital constrained.”
02 / prediction
“Next year, a big new entrant is, for example, SpaceX, which is building a ton of compute. They’re actively going to lease quite a bit of it to Anthropic and OpenAI, most likely, because they’re the ones who have the marginal capability to pay the highest price.”
03 / belief
“There was the whole spat recently where I think Gavin Baker was like, “Dario believes that there’s only going to be one company in the world.”
04 / belief
“I think most compute will still continue to transact at sub-$20 billion a gigawatt.”
05 / evaluation
“They’re not adding $25 billion of ARR every month now. That means the marginal megawatt they’re getting is going as a higher percentage to R&D than it is to inference.”
06 / evaluation
“Because they know their revenue from it’s going to be huge, and they’re going to pay 20% because it’s still better than renting it from SpaceX for $50 billion a gigawatt.”
07 / observation
“You’ve seen people do funny arbitrages here where they buy turbines and then try and resell them, because the value of a turbine is way more since it’s the thing bottlenecking your data center.”
08 / prediction
“Anthropic not releasing what their safety assessment says is Model 2, which is widely believed to be the next version of Mythos. They’re clearly not releasing their best models, in which case their revenue per megawatt stalls or can even start to decline again because other models are competitive again.”
09 / evaluation
“Many of these hyperscalers were building infrastructure without knowing if there was going to be a payoff. So ultimately you had this negative value being created on the model layer, if you will, because they were selling the tokens for less than it cost them on the infra side.”
10 / prediction
“Just because someone has raised prices doesn’t mean the entire supply chain rebalances immediately.”
11 / evaluation
“When Amazon is serving Bedrock Anthropic models, that counts as Anthropic compute in our worldview, because it is effectively, at the end of the day, counted as revenue for Anthropic even though there’s a revenue share and credit back all that.”
12 / evaluation
“We can talk all we want about how they went from $20 million per megawatt to $100 million per megawatt, but they’re still paying $13 million for a lot of the compute they’re buying. But at the end of the day, the reason they’ve gone to $100 million per megawatt is because Jane Street is capturing $300 million per megawatt or $500 million per megawatt.”
13 / evaluation
“TSMC’s margins on high-performance computing—HPC, AI chips, et cetera—are higher than they are for mobile, because they have a bigger advantage in HPC than they do in mobile.”
14 / evaluation
“4 is both way cheaper to run than GPT-4 and has fewer active parameters. It’s much smaller, in that sense of active parameter, because it’s a sparser MoE versus GPT-4 being a coarser MoE.”
15 / evaluation
“If a Hopper can make a million tokens of Opus and it can make two million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from $2 to $3.”
16 / belief
“As we get to the end of the decade, we think something like half of the capacity that’s being added will be behind the meter.”
17 / belief
“Take a gigawatt of Nvidia’s Rubin chips. Rubin is announced at GTC, I believe the week this podcast goes live.”
18 / observation
“The problem is that getting the heat out of that dense area means you have to move away from standard air and liquid cooling to more exotic forms of liquid cooling, or even immersion, to get to higher power densities.”
19 / belief
“I think going back to the earlier view that if the models are so powerful, the value of a GPU goes up over time, right now only OpenAI and Anthropic have that viewpoint.”
20 / belief
“There are a couple of companies and I think that could be a big disruption to the industry beyond EUV.”
21 / belief
“I think there are a couple factors here. ASML has not decided to just go YOLO, let’s expand capacity as fast as possible.”
22 / uncertainty
“Then you’ll have this insane margin that ASML and TSMC should have been charging. But the thing is, I don’t know if ASML and TSMC will ever agree to this.”
23 / prediction
“In 2024, we were banging on the drums that reasoning means long context, which means a large KV cache, which means you need a lot of memory demand.”
24 / prediction
“In some cases, the argument people are making is if you didn’t sign a long-term deal, because every two years NVIDIA is tripling or quadrupling the performance while only 2X-ing or 50% increasing the price… Then the price of an H100… Sure maybe the value in the market was $2 at 35% gross margins in 2024, but in 2026, when Blackwell is in super high volume and deploying millions a year, you’re actually now worth $1/hour.”
25 / evaluation
“You can look across the space at hedge funds and look at their 13Fs and see they own, maybe not exactly what Leopold does, because it’s always a question of what is the most constrained thing.”
26 / prediction
“I think space data centers will eventually be a 10X gain as Earth’s resources get more and more contentious, but that’s not this decade.”
27 / commitment
“We’ll purposely undershoot what we think we can possibly do and be conservative because we don’t want to potentially go bankrupt.”
28 / prediction
“Even if you halve smartphone volumes, because of the shape of the halving, the low end gets cut by more than half, while the high end gets cut by less than half, because you and I will still buy the high-end phones that cost north of a thousand dollars.”
29 / evaluation
“If you look at PJM, which I think is the largest grid in America—covering the Midwest and some of the Northeast area—in their models they want to have roughly 20 percent excess capacity.”
30 / evaluation
“Google even had to go to TSMC and explain to them why they needed this increase in capacity because it was so sudden.”
31 / evaluation
“There’s a lot of R&D, there’s a lot of customer acquisition costs. This is sort of why, not Microsoft, but the SaaS companies have underperformed massively in the markets, because the COGS of AI is just so high, and that just completely breaks how these business models work.”
32 / evaluation
“You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs because now if you think about the parameter count as the total embedding space for all of this knowledge that you’re compressing down during training, one, you’re embedding this data in instead of having to activate every single parameter, every single time you’re training or running inference, now you can just activate on a subset and the model will learn which expert to route to for different tasks.”
33 / belief
“I think actually Google’s interface is sometimes nice, but it’s also they don’t care about anyone besides their top customers.”
34 / belief
“All of these things can go happen faster. And so I think software and then the other domain is industrial, chemical, mechanical engineers suck at coding just generally.”
35 / belief
“I think four of Amazon’s top five revenue products, margin products like gross profit products are all database-related products like Redshift and all these things.”
36 / belief
“I think everyone has benefited regardless because the data’s on the internet. And therefore, it’s in your per training now.”
37 / belief
“I think there’s a couple factors here. One is that they do have model architecture innovations.”
38 / belief
“I think to some extent, we have capabilities that hit a certain point where any one person could say, “Oh, okay, if I can leverage those capabilities for X amount of time, this is AGI, call it ’27, ’28.”
39 / belief
“I would say there’s two articles in this one that I could show maybe graphics that might be interesting for you to pull up.”
40 / evaluation
“What Nathan’s referring to is in 2020, Huawei released their Ascend 910 chip, which was an AI chip, first one on seven nanometer before Google did, before NVIDIA did. And they submitted it to the MLPerf benchmark, which is sort of a industry standard for machine learning performance benchmark, and it did quite well, and it was the best chip at the submission.”
41 / evaluation
“Not by a lot though, right? And R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas.”
42 / belief
“I think the unsung heroes are the cooling in electrical systems which are just glossed over.”
43 / uncertainty
“I think OpenAI’s statement, I don’t know if you’ve seen the five levels where it’s chat is level one, reasoning is level two, and then agents is level three.”
44 / belief
“China will win because of these restrictions long-term, unless AI does something in the short-term, which I believe AI will make massive changes to society in the medium, short-term.”
45 / belief
“Better lesson, right? The question is, I think, when not if, because the rate of progress is so fast.”
46 / uncertainty
“We don’t know what it is, just that they are not just issuing one chain of thought in sequence.”
47 / evaluation
“YOLO.” And this is where that sort of stress comes in is like, “Well, I know it works here, but some things that work here don’t work here.”
48 / belief
“I think if you look at every layer of the compute stack, whether it goes from lithography and etch all the way to fabrication, to optics, to networking, to power, to transformers, to cooling, to a networking, and you just go on up and up and up and up the stack, even air conditioners for data centers are innovating.”
49 / belief
“I think, and there were a lot of false narratives, which is like, “Hey, these guys are spending billions on models,” and they’re not spending billions on models.”
50 / belief
“I think one of the few things you can purchase ironically, is a Texas Instruments graphing calculator because they actually manufacture in Texas.”
51 / belief
“I think the thing that’s really important about these mega cluster buildouts is they’re completely unprecedented in scale.”
52 / belief
“Of what GenAI can do to society, but it was very clear, I think, to at least National Security Council and those sort of folks, that this was where the world is headed, this cold war that’s happening.”
53 / belief
“I think there is one aspect to note though is that there is the general ability for that to transfer across different types of runs.”
54 / belief
“I think others will keep raising money because the returns from it are going to be eventually huge once we have AGI.”
55 / belief
“I think when you look at the Chinese labs, Huawei has a lab, Moonshot AI, there’s a couple other labs out there that are really close with the government, and then there’s labs like Alibaba and DeepSeek, which are not close with the government.”
56 / belief
“Generally, humans have positive impacts on the world, at least societally, but it’s possible for individual humans to have such negative impacts. And AGI, at least as I think the labs define it, which is not a runaway sentient thing, but rather just something that can do a lot of tasks really efficiently amplifies the capabilities of someone causing extreme damage.”
57 / recommendation
“Maybe we use some sort of reward model outside of this to select even the best one to preference, as well.”
58 / evaluation
“The beautiful thing about AI is that because it's growing so fast, every layer is being stressed to an incredible degree.”
59 / belief
“You should tell that story, actually, about the TSMC guy that went to Samsung and SMIC and all that. I think you should tell that story.”
60 / observation
“The way to look at it today is that it’s super stratified. Every industry has anywhere from one to three competitors.”
61 / uncertainty
“What we see beyond that is more questionable and I'm not sure because I don't know.”
62 / belief
“I think more reasonably, it will be like the total summation of the flops that you deliver to the model across pre-training, post-training, synthetic data for that pre-training and post-training data, as well as some of the inference time compute efficiencies.”
63 / belief
“You have multiple, at least five, I believe GB200 100K clusters being built by Microsoft/OpenAItheir partners for them.”
64 / belief
“You could just destroy every GPU in a data center if you want if you just fuck with the grid, pretty easily, I think.”
65 / belief
“Next year, 300-500k depending on whether it's one site or many. 300-700k I think is the upper bound of that.”
66 / belief
“We think, with fairly high accuracy, that there are five regions that they're connecting together, which comprises many data centers.”
67 / belief
“I think Jon and I could communicate to the point where you even wouldn't know what we're talking about.”
68 / evaluation
“All in all, you only save about 20% in power per transistor. But because of data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node than just the individual transistors' power efficiency benefit.”
69 / preference
“You just go down the list. That means no fridges, no automobiles, no weed whackers, because that stuff has chips.”
70 / prediction
“Anyway, there's a tremendous opportunity to bring breakthrough innovation simply because there are so many layers where things are unoptimized.”
71 / prediction
“I'd bet if AI is as important as you and I believe, they will centralize sooner than the West does.”
72 / observation
“What has happened is the output per individual has soared because of EDA (Electronic Design Assistance) tooling.”
73 / recommendation
“If you want to restrict them from having chips, you have to let them have at least some level of chip that is better than what they can build internally.”
74 / evaluation
“Then as soon as I started making money… I grew up in a family business, but I didn't get paid for working, of course, like yourself.”
75 / uncertainty
“I don't know if we could tell because they could also just easily spawn like 10 other aluminum mills to make up for the production and be fine.”
76 / prediction
“China has no issue with that at all because their supply chain adds as much power as like half of Europe every year.”
77 / preference
“” The way I like to look at it is that chip manufacturing is like 3D chess or like a massive jigsaw puzzle.”
78 / uncertainty
“I mean, I don't know if I, if I have a coherent thesis, but it's, it's sure fun to, it's Who, who think that like, I, I, I just have an intense hatred for RAG.”
79 / evaluation
“can ballpark things like, you know, like, yeah, so, so, so, but basically like the, the point is like you're trading this is optimal, this is theoretically the most optimal architecture for performance power and area in a given, and you know, not, not specifically Grok, but VLIW in general is gonna get you closer to optimal there, but then you're giving off, you know, that, that last P, which is pain in the ass program, is, is I think the most simple way to get into it.”
80 / belief
“Like, you look at what Tsinghua University is doing in China, actually, they open sourced their model to I think the largest like by parameter count, at least open source models.”
81 / belief
“I think with AI, especially like how scaling laws are going, it's like incredibly important for infrastructure is like so much more important. And then like when you just think about software cost, right, like the cost structure of it, there was always a bigger component of R&D and like SAS businesses, you know, all over SF, all these SAS businesses did crazy good because, you know, they just start as they grow and then all of a sudden they're so freaking profitable for each incremental new customer.”
82 / evaluation
“Building a bigger and better model every, you know, every few months. And I don't know how Apple gets on that train, but you know, at the same time, there's no company that has more powerful distribution, right?”
83 / prediction
“I don't even think like, I expect in the next few years that the OpenAI Microsoft probably falls apart too.”
84 / evaluation
“The supercomputer, it's, it's, oh, it's slightly more, but yeah, I think the 500 million is a fair enough number.”
85 / evaluation
“Their chips are not as good, or if they are, even though, you know, I mentioned Intel and AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?”
86 / evaluation
“I mean, it was good. They were doing good, and NVIDIA bought them, you know, in 19, I believe, or 18, but Broadcom has been number one in networking for a decade plus, and Google partnered with them on making the TPU, right?”
87 / recommendation
“Unless you're fine tuning for on-device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time because the cost of inference is actually much cheaper than you think once you achieve good MBU and you batch at a decent size, which any successful business in the cloud is going to achieve, you know, and then two, fine tuning like people like, oh, you know, this 7 billion parameter model, if you fine tune it on a data set is almost as good as 3.”
88 / preference
“Some people would argue lower, right? But at the very least, you need to achieve human reading level speeds and probably a little bit faster, because we like their skin, to have a usable model for chatbot style applications.”
89 / evaluation
“I think that one's really critical because it explains Google's infrastructure quite a bit from networking through chips, through all that sort of history of the TPU a little bit.”
90 / prediction
“So in training, everyone just talks about MFU, right? But then on inference, right, which I think is one LLM inference will be bigger than training or multimodal, whatever, bubble inference will be bigger than training, probably next year, in fact, at least in terms of GPUs deployed.”