The industry panicked, you know, that this is the end of AI. And of course, of course that, that's obviously not true. We're gonna keep on scaling the amount of data that we have to, to train with.
Jensen Huang
Recommendations and personal stack
inference is gonna be the biggest market, and it's gonna be easy, and we're gonna commoditize it. You know, everybody can build their own chips.
And thank you for releasing open source Nemotron 3 Super, which you can also use inside Perplexity to look stuff up. Now, which is 120 billion parameter open weight MoE model.
And thank you for releasing open source Nemotron 3 Super, which you can also use inside Perplexity to look stuff up. Now, which is 120 billion parameter open weight MoE model.
One of the things that I love about Nemotron 3 is it's not a, just a pure transformer model, it's transformer and SSMs.
By the way, GeForce is our still, to this day- ... our number one marketing strategy. Right. People learn about NVIDIA while they're in their teenage years.
By the way, GeForce is our still, to this day- ... our number one marketing strategy. Right. People learn about NVIDIA while they're in their teenage years.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
It really brought a lot of joy to a lot of people. The, the, the hardware really brings these worlds to life.
Published claims
And so we just gotta bring every technology to bear. Otherwise, we scale up linearly or we scale up based on the capabilities of Moore's Law, which has largely slowed because Dennard scaling has slowed.
So, CUDA turned out to be an incredible foundation for computation in this AI infrastructure world. So-
you know, you know how it is. You manifest a future and that future is so convincing, there's no way it won't happen. There's a lot of suffering in between, but you've gotta believe what you believe.
The industry panicked, you know, that this is the end of AI. And of course, of course that, that's obviously not true. We're gonna keep on scaling the amount of data that we have to, to train with.
So I think you've outlined four of them with pre-training, post-training, test time, and agentic scaling.
inference is gonna be the biggest market, and it's gonna be easy, and we're gonna commoditize it. You know, everybody can build their own chips.
These AI model architectures are being invented about once every six months. Right? And system architectures and hardware architectures kind of every three years. And so you need to anticipate what likely is going to happen, you know, two, three years from now.
That's ridiculous. Let's use the... Let's use a thought experiment. And you could just sit there, enjoy a glass of whiskey, and, and think about all these things, and it would become completely obvious. Like, if I were to create the most amazing- the most amazing agent that we can imagine in the next 10 years. Let's say it'd be a humanoid robot. If that humanoid robot were to be created, is it more likely that the humanoid robot comes into my house and uses the tools that I have to do the work that it needs to do? Or does this hand turns into a 10-pound hammer in one instance, turns into a scalpel in another instance, and in order to boil water, it beams, you know, microwaves out of its fingers?
If you take the OpenClaw schematic that I used at GTC, you'll find it two years ago. Literally, two years ago at GTC, I was talking about agentic systems that exactly reflect OpenClaw today.
There's really serious and complicated security concerns about when you have such powerful technology, how do you hand over your data so they can do useful stuff? But then there's scary things associated with that. And we, as a civilization, as individual people and as a civilization, figuring out how to find that right balance.
he made it his... He makes it his business that he's the top priority of everybody else's, you know, projects. And so he does that by demonstrating it.
he literally is going through the entire process of how to plug in cables into a rack. He's, he's working with an engineer on the ground that's doing that task, and he's just trying to understand what does that process look like so it can be less error-prone.
We open sourced the models, we open sourced the weights, we open sourced the data, we open sourced how we created it. Yeah, it's pretty amazing.
And so it just depends on somehow they've, they've balanced these two and they're world-class at both.
We're the largest computer company in history. That alone should beg the question, why?
The answer is of course yes. And the reason for that is because it's not limited by any physical limits. There's nothing that I see that says, you know, gosh, $3 trillion is not possible.
Is there something truly, as you know, something truly special happening from about December, where people have really woke up to the power of Claude Code of Codex, of OpenClaw.
NVLink-72 literally builds supercomputers in the supply chain and ships 'em two, three tons at a time per rack.
It used to be—they used to come in parts and we used to assemble 'em inside the data center. But that's impossible now because NVLink-72 is so dense.
The market for inference is, you know, coming. The inflection point for inference is coming. It's gonna be a big market.
I first explain to them what's going on, why it's gonna happen, and then I ask 'em to make several billion dollars of capital investments each. And because they, you know, they trust me and I'm very respectful of 'em, and I give 'em every opportunity to question me and I spend time to explain things to people and I reason about it. I draw on pictures and I reason about it in first principles.
it's a lot of it is about relationships and building a shared view of the future.
our power grid is designed for the worst case condition with some margin. Well, 99% of the time we're nowhere near the worst case condition because the worst case condition is a few days in the winter, a few days in the summer, and extreme weather. Most of the time we're nowhere near the worst case condition and we're probably running around, call it 60% of peak.
99% of the time, our power grid has excess power, and they're just sitting idle, but they have to be there sitting idle because just in case, when the time comes, hospitals have to be powered and, you know, infrastructure has to be powered and airports have to run and so on and so forth.
we could go and, Help them understand and create contractual agreements and design computer architecture systems, data centers, such that when they need, The maximum power for infrastructure in society, that the data centers would get less. But that's in a very rare instance anyways. And during that time, we either have a backup generator for that little part of it, or we just have our computers shift the workload somewhere else, or we have the computers just run slower. You know, we could degrade our performance, reduce our power consumption and provide for a, you know, slightly longer latency response
it's putting a lot of pressure on the grid to be able to... Now, they're gonna have to increase from their maximum. I just wanna use their excess. It's just sitting there.
Our country's leaders, incredible, but they're mostly lawyers. Their country's leaders, and because we're, they're trying to keep us safe, rule of law, governing, their country was built out of poverty. And so most of their leaders are incredible engineers. Some of the brightest minds.
And thank you for releasing open source Nemotron 3 Super, which you can also use inside Perplexity to look stuff up. Now, which is 120 billion parameter open weight MoE model.
First, if we're gonna be a great AI computing company, we have to understand how AI models are evolving.
One of the things that I love about Nemotron 3 is it's not a, just a pure transformer model, it's transformer and SSMs.
And we were early in, Developing the conditional GANs, which, that progressive GANs, which led step-by-step to diffusion.
The fact that we're doing basic research in model architecture and in different domains gives us visibility into, you know, what kind of computing systems would do a good job for future models.
Second, um, I think we rightfully recognize that on the one hand, we want world-class models as products, and they should be proprietary. On the other hand, we also want AI to diffuse into every industry and every country, every researcher, every student.
Open source is fundamentally necessary for many industries to join the AI revolution.
NVIDIA has the scale and we have the motives to not only skills, scale, and motivation to build and continue to build these AI models for as long as we shall live.
These AIs will likely use tools and models and sub-agents that were trained on other modalities of information. Maybe it's biology or chemistry or you know, laws of physics, or you know, fluids and thermodynamics, and not all of it is in language structure.
if you knew how hard it would be to build NVIDIA it turned out to be, what is it? A million times more hard than you anticipated that you wouldn't do it.
you need to have endurance, you need to have grit, so that when the setbacks actually happened, and those setbacks are gonna surprise you, the disappointments are gonna surprise you, you know, the embarrassments are gonna surprise you, the humiliations are gonna surprise you.
I'm constantly reasoning through, "Let me tell you what, how I see it." And then I reason through it. It gives everybody the opportunity to intercept and say, "I disagree with that part." The nice thing about reasoning through things and letting people interact with it is that they don't have to disagree with your outcome. They can disagree with your reasoning steps.
And so we're kind of, you know, collective path searching method. And it's really fantastic.
They knew that recently my first job was, you know, cleaning toilets, so.
By the way, GeForce is our still, to this day- ... our number one marketing strategy. Right. People learn about NVIDIA while they're in their teenage years.
And then they go to college and they know who NVIDIA is and then in beginning is just, you know, playing Call of Duty, you know? You know, Fortnite. And then later they're using CUDA, and then later they're using NVIDIA and, you know, Blender and Dassault and Autodesk.
It really brought a lot of joy to a lot of people. The, the, the hardware really brings these worlds to life.
study scans at so much faster now, you could study more scans, you could diagnose better, you could, you could inpatient faster, you can see people more. The hospitals are making more money. You have more patients in the hospital. You need more radiologists.
The number of software engineers at NVIDIA is gonna grow, not decline. And the reason for that is because the purpose of a software engineer and the task of a software engineer coding are related, not the same. I wanted my software engineers to solve problems.
the people that are currently programmers and software engineers, I think they're at the cutting edge of understanding intuitively how to communicate with the agents using natural language in order to design the best kind of software.
the goal of specification, the artistry of specification, the goal and the artistry of it, Is going to depend on what problem you're trying to solve.
I believe that AI will be able to recognize those and understand those.
I don't think my chips will feel those. And therefore, the... How that anxiety, how that feeling, how that excitement, how that, how that, you know...
I don't think there's anything about anything that we're building that would suggest that two different computers being presented with all of exactly the same context would perfo- Of course, it would produce statistically different outcomes, but it's not because it felt different.
Scaling can create some incredible miracles in the space of intelligence. Has been truly marvelous to watch, so I'm open to surprise.
Intelligence has a meaning, you know? And it's a system that... You know, it's something that we do that includes perception and understanding and reasoning and the ability to do plan.
Intelligence is not one word that is exactly equal to humanity. And that's, I think it's really important to separate the two.
I think AI will help us celebrate humans more. And certainly humanity and human first, and I, I think what makes this world incredible is humans forever will be so, and just AI is this incredible tool that makes us-
thank you for all the interviews that you do, the depth, the respect that you go through with and the research that you do to reveal, you know, for all of us, The amazing people that you've interviewed over the years.
as an innovator, to have created this long form, unbelievable, and yet, you know, it's just captivating.