Evidence receipt / prediction
Published · transcript-backedJensen Huang: prediction
15 Apr 2026 Dwarkesh Podcast Jensen Huang – TPU competition, why we should sell chips to China, & Nvidia’s supply chain moat
“The only way to really get 10x or 100x leaps is to fundamentally change the algorithm and how it’s computed every single year.”
Source trail
Everything needed to verify it.
- Speaker
- Jensen Huang
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 15 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…This is going to be a long question. You have spectacular revenue, and you’re not making $60 billion a quarter from pharma and quantum. You’re making it because AI is an unprecedented technology that is growing unprecedentedly fast. The question then is what is best for AI specifically. I’m not in the details, but I talk to my AI researcher friends and they say, “Look, when I use a TPU, it’s this big systolic array that’s perfect for doing matrix multiplies, whereas a GPU is very flexible. It’s great when you have lots of branching or irregular memory access.” But what is AI? It’s just these very predictable matrix multiplies again and again and again. You don’t have to give up any die area for warp schedulers or switches between threads and memory banks. And the TPU is really optimized for the bulk of this growth in revenue and use case for compute that is coming online right now. I wonder how you react to that. Matrix multiplies are an important part of AI, but they’re not the only part. If you want to come up with a new attention mechanism, disaggregate in a different way, or invent a whole new type of architecture altogether—like a hybrid SSM—you want an architecture that’s generally programmable. If you want to create a model that fuses diffusion and autoregressive techniques, you want an architecture that’s just generally programmable. We run everything you can imagine. That’s the advantage. It allows for the invention of new algorithms a lot more easily, because it’s a programmable system. The ability to invent new algorithms is really what makes AI advance so quickly. TPUs, like anything else, are impacted by Moore’s Law, which we know is increasing by about 25% per year. The only way to really get 10x or 100x leaps is to fundamentally change the algorithm and how it’s computed every single year. That’s Nvidia’s fundamental advantage. The only reason we were able to make Blackwell to Hopper 50x… When I first announced Blackwell was going to be 35x more energy efficient than Hopper, nobody believed it. Then Dylan wrote an article saying I sandbagged, and it’s actually fifty times. You can’t reasonably do that with just Moore’s Law. The way we solve that problem is with new models, like MoEs, that are parallelized, disaggregated, and distributed across a computing system. Without the ability to really get down and come up with new kernels with CUDA, it’s really hard to do. It’s the combination of the programmability of our architecture and the fact that Nvidia is an extreme co-design company. We can even offload some of the computation into the fabric itself, like NVLink, or into the network with Spectrum-X. We could affect change across the processors, the system, the fabric, the libraries, and the algorithm simultaneously. Without CUDA to do that, I wouldn’t even know where to start. This gets at an interesting question about Nvidia’s clientele. 60% of your revenue is coming from these big five hyperscalers. In a different era with different customers—let’s say professors running experiments—they need CUDA. They can’t use another accelerator. They just needed to run PyTorch with CUDA and have everything optimized. But these hyperscalers have the resources to write their own kernels. In fact, they have to in order to get that last 5% of performance they need for their specific architecture. Anthropic and Google are mostly running their own accelerators or running TPUs and Trainium. But even OpenAI, using GPUs, has Triton because they need their own kernels. Down to CUDA C++, instead of using cuBLAS and NCCL, they’ve got their own stack which compiles to other accelerators as well. If most of your customers can and do make replacements for CUDA, to what extent is CUDA really the thing that is going to make frontier AI happen on Nvidia?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.