Evidence receipt / evaluation
Published · transcript-backedJensen Huang: evaluation
15 Apr 2026 Dwarkesh Podcast Jensen Huang – TPU competition, why we should sell chips to China, & Nvidia’s supply chain moat
“With post-training and reinforcement learning, that entire area is just exploding. So if you want to build on an architecture, building on CUDA makes the most sense because you know the ecosystem is great.”
Source trail
Everything needed to verify it.
- Speaker
- Jensen Huang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…This gets at an interesting question about Nvidia’s clientele. 60% of your revenue is coming from these big five hyperscalers. In a different era with different customers—let’s say professors running experiments—they need CUDA. They can’t use another accelerator. They just needed to run PyTorch with CUDA and have everything optimized. But these hyperscalers have the resources to write their own kernels. In fact, they have to in order to get that last 5% of performance they need for their specific architecture. Anthropic and Google are mostly running their own accelerators or running TPUs and Trainium. But even OpenAI, using GPUs, has Triton because they need their own kernels. Down to CUDA C++, instead of using cuBLAS and NCCL, they’ve got their own stack which compiles to other accelerators as well. If most of your customers can and do make replacements for CUDA, to what extent is CUDA really the thing that is going to make frontier AI happen on Nvidia? CUDA is a rich ecosystem. If you want to build on any computer first, building on CUDA first is incredibly smart. Because the ecosystem is so rich, we support every framework. If you want to create custom kernels… For example, we contribute enormously to Triton. So the back end of Triton has huge amounts of Nvidia technology. We’re delighted to help every framework become as great as it can be. There are lots and lots of frameworks. There’s Triton, vLLM, SGLang, and more. Now there’s a whole bunch of new reinforcement learning frameworks coming out, like verl and NeMo RL. With post-training and reinforcement learning, that entire area is just exploding. So if you want to build on an architecture, building on CUDA makes the most sense because you know the ecosystem is great. You know that if something happens, it’s more likely in your code and not in the mountain of code underneath. Don’t forget the amount of code you’re dealing with when building these systems. When something doesn’t work, was it you or was it the computer? You would like it to always be you and to be able to trust the computer. Obviously, we still have lots of bugs ourselves, but our system is so well wrung out that you can at least build on top of the foundation. That’s number one: the richness, programmability, and capability of the ecosystem. The second thing is, if you’re a developer building anything at all, the single most important thing you want is an install base. You want the software you write to run on a whole bunch of other computers. You’re not building software just for yourself. You’re building it for your fleet or everybody else’s fleet because you’re a framework builder. Nvidia’s CUDA ecosystem is ultimately its great treasure. We have several hundred million GPUs out there now. Every cloud has it. It goes back to the A10, A100, H100, H200, the L series, the P series. There’s a whole bunch of them. They’re in all kinds of sizes and shapes. If you’re a robotics company, you want that CUDA stack to actually run in the robot itself. We’re literally everywhere. The install base means that once you develop the software or the model, it’s going to be useful everywhere. That is just incredibly valuable. Lastly, the fact that we’re in every single cloud makes us genuinely unique. If you’re an AI company or developer, you’re not exactly sure which cloud service provider you’re going to partner with or where you’d like to run it. We run everywhere, including on-prem for you if you like. The combination of the richness of the ecosystem, the expansiveness of the install base, and the versatility of where we are makes CUDA invaluable. That makes a lot of sense. I guess the thing I’m curious about is whether those advantages matter a lot to your main customers. There’s many people for whom they might matter. The kind of person who can actually build their own software stack makes up most of your revenue. Especially if you go to a world where AI is getting especially good at the things which have tight verification loops where you can RL on them…. This question of how do you write a kernel that does attention or MLP the most efficiently across a scale up? It’s a very verifiable sort of feedback loop. Can all the hyperscalers write these custom kernels for themselves? Nvidia still has great price performance, so they might still prefer to use Nvidia. But then the question is, does it just become a question of who is offering the best specs, the best flops and memory bandwidth for a given dollar. Whereas historically Nvidia has just had, and still has, the best margins in all of AI across hardware and software, +70%, because of this CUDA moat. And the question is, can you sustain those margins if for most of your customers, they can actually afford to build, instead of the CUDA moat?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.