Evidence receipt / evaluation
Published · transcript-backedJensen Huang: evaluation
23 Mar 2026 Lex Fridman Podcast #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
“You know, everybody can build their own chips. And, and that was always illogical to me because inference is thinking, and I think thinking is hard.”
Source trail
Everything needed to verify it.
- Speaker
- Jensen Huang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Mar 2026
- Publisher
- Lex Fridman Podcast
Transcript context
…So I think you’ve outlined four of them with pre-training, post-training, test time, and agentic scaling. What do you think, when you think about the future, deep future and the near-term future, what are the blockers that you’re most concerned about that keep you up at night that you have to overcome in order to keep scaling? Well, we can go back and reflect on what people thought were blockers. So in the beginning, we were… The pre-training scaling law. You know, people thought, rightfully so, that the amount of data that we have, high-quality data that we have, will limit the intelligence that we achieve. And that scaling law was an important, very important scaling law. The larger the model, the correspondingly more data results in a smarter AI. And so that was pre-training. And Ilya Sutskever said, “We’re out of data,” or something like that. “Pre-training is over,” or something like that. The industry panicked, you know, that this is the end of AI. And of course, that’s obviously not true. We’re gonna keep on scaling the amount of data that we have to train with. A lot of that data is probably gonna be synthetic, and that also confused people, you know? And what people don’t realize is they’ve kind of forgotten that most of the data that we are training, that we teach each other with, inform each other with, is synthetic. You know, it’s synthetic because it didn’t come out of nature. You created it. I’m consuming it. I modify it, augment it, I regenerate it, somebody else consumes it. And so we’ve now reached a level where AI is able to take ground truth, augment it… Enhance it, synthetically generate an enormous amount of data. And that part of post-training continues to scale, and so the amount of data that we could use that is human generated will be smaller, and smaller, and smaller. The amount of data that we use to train models is going to continue to scale to the point where we’re no longer limited… Training is no longer limited by… Data is now limited by compute. And the reason for that is most of the data is synthetic. Then the next phase is test time, and I still remember people telling me that, “Inference? Oh, yeah, that’s easy. Pre-training, that’s hard.” These are giant systems that people are talking about. Inference must be easy. And so inference chips are gonna be little tiny chips, and- … you know, they’re not, they’re not like NVIDIA’s chips. Oh, those are gonna be complicated and expensive, and, you know, we could make… And this is- … in, in the future, inference is gonna be the biggest market, and it’s gonna be easy, and we’re gonna commoditize it. You know, everybody can build their own chips. And, and that was always illogical to me because inference is thinking, and I think thinking is hard. Thinking is way harder than reading. ommoditize it. You know, everybody can build their own chips. And, and that was always illogical to me because inference is thinking, and I think thinking is hard. Thinking is way harder than reading. You know, pre-training is just memorization and generalization, you know, and looking for patterns in relationships. You’re reading and reading, versus thinking, reasoning, solving problems, taking unexplored experiences, new experiences, and breaking it down into… Decomposing it into, you know, solvable pieces that we then go off, either through first principle reasoning, or, you know, through previous examples, prior experiences. You know, or just exploration and search and, you know, trying different things. And that whole process of test time scaling inference, is really about thinking. And it’s about reasoning, it’s about planning, it’s about search, it’s about… And so how could that possibly be compute light? And we were absolutely right about that. You know, so test time scaling is intensely compute intensive. Then the question is, okay, now we’re at inference and we’re at test time scaling, what’s beyond that? Well, obviously we have now created, you know, one agentic person, and that one agentic person has a large language model that we’ve now developed. But during test time, that agentic system goes off and does research and bangs on databases, and it goes out and, you know, uses tools, and one of the most important things it does is spins off and spawns off a whole bunch of sub-agents. Which means we’re now creating large teams. It’s so much easier to scale NVIDIA by hiring more employees than it is to scale myself. And so the next scaling law is the agentic scaling law. It’s kind of like multiplying AI. Multiplying AI, we could spin off agents as fast as you want to spin off agents. And so, you know, I… You know, I have four scaling laws. And as we use the agentic systems, they’re gonna create a lot more data, they’re gonna create a lot of experiences. Some of it we’re gonna say, “Wow, this is really good. We ought to memorize this.” That data set then comes all the way back to pre-training. We memorize and generalize it. We then refine it and fine-tune it back into post-training. Then we enhance it even more with test time, you know, and the agentic systems, you know, put it out to the industry. And so this loop, this cycle, is gonna go on and on and on. It kinda comes down to basically intelligence is gonna scale by one thing, and that’s compute.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.