Evidence receipt / belief
Published · transcript-backedSatya Nadella: belief
19 Feb 2025 Dwarkesh Podcast Satya Nadella — Microsoft’s AGI plan & quantum breakthrough
“The way I think about it is hey, distributed computing will remain distributed, so go build out your fleet such that it's ready for large training jobs, it's ready for test-time compute, it’s ready- in fact, if this RL thing that might happens, you build one large model, and then after that, there’s tons of RL going on.”
Source trail
Everything needed to verify it.
- Speaker
- Satya Nadella
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 19 Feb 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…On the hyperscaler side, and by the way, it's also interesting the advantage you as a hyperscaler would have in the sense that, especially with inference time scaling and if that's involved in training future models, you can amortize your data centers and GPUs, not only for the training, but then use them again for inference. I'm curious what kind of hyperscaler you consider Microsoft and Azure to be. Is it on the pre-training side? Is it on providing the O3-type inference? Or are you just, we’re going to host and deploy any single model that's out there in the market, and we are sort of agnostic about that? It’s a good point. The way we want to build out the fleet is [to], in some sense ride Moore's law. I think this will be like what we've done with everything else in the past: every year keep refreshing the fleet, you depreciate it over whatever the lifetime value of these things are, and then get very very good at the placement of the fleet such that you can run different jobs at it with high utilization. Sometimes there are very big training jobs that need to have highly concentrated peak flops that are provisioned to it that also need to cohere. That's great. We should have enough data center footprint to be able to give that. But at the end of the day, these are all becoming so big, even in terms of if you take pre-training scale, if it needs to keep going, even pre-training scale at some point has to cross data center boundaries. It's all more or less there. So, great, once you start crossing pre-training data center boundaries, is it that different than anything else? The way I think about it is hey, distributed computing will remain distributed, so go build out your fleet such that it's ready for large training jobs, it's ready for test-time compute, it’s ready- in fact, if this RL thing that might happens, you build one large model, and then after that, there’s tons of RL going on. To me, it's kind of like more training flops, because you want to create these highly specialized, distilled models for different tasks. So you want that fleet, and then the serving needs. At the end of the day, speed of light is speed of light, so you can't have one data center in Texas and say, "I'm going to serve the world from there." You've got to serve the world based on having an inference fleet everywhere in the world. That's how I think of our build-out of a true hyperscale fleet. Oh, and by the way, I want my storage and compute also close to all of these things, because it's not just AI accelerators that are stateless. My training data itself needs storage, and then I want to be able to multiplex multiple training jobs, I want to be able to then have memory, I want to be able to have these environments in which these agents can go execute programs. That's kind of how I think about it. You recently reported that your yearly revenue from AI is $13 billion. But if you look at your year-on-year growth on that, in like four years, it'll be 10x that. You'll have $130 billion in revenue from AI, if the trend continues. If it does, what do you anticipate doing with all that intelligence, this industrial scale use? Is it going to be through Office? Is it going to be you deploying it for others to host? You've got to have the AGIs to have $130 billion in revenue? What does it look like?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.