High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Jensen Huang: preference

15 Apr 2026 Dwarkesh Podcast Jensen Huang – TPU competition, why we should sell chips to China, & Nvidia’s supply chain moat

“Because the customers make so much money—for example, our software engineers—if I can give them much more responsive tokens so that they’re even more productive than they are today, I would pay for it.”

— Jensen Huang

Source trail

Everything needed to verify it.

Speaker
Jensen Huang
Attribution
Verified speaker
Claim type
preference
Recorded
15 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…One question somebody I was talking to had is, why doesn’t Nvidia run multiple different chip projects at the same time with totally different architecture? So you could do something like a Cerebras-style wafer scale. You could do a Dojo-style huge package. You could do one without CUDA. You have the resources and the engineering talent to do all of these in parallel. So why put all the eggs in one basket, given who knows where AI might go and architectures might go? Oh, we could. It’s just that we don’t have a better idea. We could do all of those things. It’s just not better. We simulate it all in our simulator, proveably worse. So we wouldn’t do it. We’re working on exactly the projects that we want to work on. If the workload were to change dramatically—and I don’t mean the algorithms, I actually mean the workload, and that depends on the shape of the market—we may decide to add other accelerators. For example, recently we added Groq, and we’re going to fold Groq into our CUDA ecosystem. We’re doing that now because the value of tokens has gone up so high that you could have different pricing of tokens. Back in the old days, just a couple years ago, tokens were either free or barely expensive. But now you can have different customers, and those customers want different answers. Because the customers make so much money—for example, our software engineers—if I can give them much more responsive tokens so that they’re even more productive than they are today, I would pay for it. But that market has only recently emerged. So I think we now have the ability to have the same model, based on the response time, have different segments. That’s the reason why we decided to expand the Pareto frontier and create a segment of inference that is faster response time, even though it’s lower throughput. Until now, higher throughput is always better. We think there could be a world where there could be very high ASP tokens, and even though the throughput is lower in the factory, the ASPs make up for it. That’s the reason why we did it. But otherwise, from an architecture perspective, if I had more money, I would put more behind Nvidia’s architecture. I think this idea of extremely premium tokens and just the disaggregation of the inference market is a very interesting.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence