Evidence receipt / observation
Published · transcript-backedSpeaker unverified: observation
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“I mean at this you really of course want to continue to get improved improvements, peak improvements on HPM bandwidth and chip to chip communication by using the latest technology. But then the key is how effectively do you use that bandwidth, how effectively do you use that communication and do make sure that you don't waste it, right.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- observation
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah so Samba Nova was was founded 2017 and it was kind of a out of ideas from Chris Ray my Co founder is also a professor at Stanford and certified genius. And the idea was to, you know, if you could bring software algorithm ideas together with hardware architecture ideas. And as you said in your introduction, I've been working in the in the hardware architecture space for a long time. And you know, starting from a clean slate, how would you design an architecture that's optimized specifically for inference, right? So you know, everybody thinks about GPU's as a kind of general purpose computing substrate, right? But you know, there originally was designed for graphics and then they kind of made a foray into kind of high performance computing. And for high, high performance computing, of course you need a lot of matrix calculation capability. And at some point people realise that hey, you know, you could use these things for executing machine language of machine learning models. And the core of machine learning and of course AI is matrix multiplication, right? But when you want to train a model, clearly the core of the problem is how quickly you can do very, very large matrix multiplications. And So what happens is over time, GPU's put more and more of their silicon area into the exercise of, of making these matrix multiplication capabilities better using the tensor cores. But once you train the model, right, and you train the model once, you now need to use that model, of course. And that's the inference problem. And the inference problem is not really a compute problem because as the models get bigger, you now need to move the weights and of course, what we call the KV cache into the compute units. And that is essentially a data movement problem, right? And it's a data movement problem from the memory to the compute units. And it's a data movement problem from, you know, your chip compute unit. And of course you need to scale to multiple chips in order to handle the computational requirements for a very especially for very low latency, high speed inference. And so our focus was how you design an architecture that minimises the overhead of compute of communication and make sure that you can most efficiently use the core resource in the in the system, which is the memory. And it's the memory. Memory isn't just one thing as you all know, it's a hierarchy of memories, right? And so the key thing is how do you orchestrate the hierarchy, how you orchestrate the communication such that you keep everything as efficiently utilized as possible. And if you do it right, you can get A5 to 10X improvement over where where GPU's are today. It strikes me that Nvidia's kind of solution around this has just been to increase the bandwidth rather than with through NB Link and advanced HBM integration and software optimizations like Tensor MTBLM. So is it, are they actually trying to brute force their way into this? Yeah. e the bandwidth rather than with through NB Link and advanced HBM integration and software optimizations like Tensor MTBLM. So is it, are they actually trying to brute force their way into this? Yeah. I mean at this you really of course want to continue to get improved improvements, peak improvements on HPM bandwidth and chip to chip communication by using the latest technology. But then the key is how effectively do you use that bandwidth, how effectively do you use that communication and do make sure that you don't waste it, right. And so whereas GPU's are often running at maybe 10 to 20% of the the capabilities of the resources, right, the bandwidth and, and the the memory bandwidth and the communication resources, our goal in a Salmonova system is to push that to be 70 to 80% of the peak. And so the idea is, yeah, everybody wants more capabilities from the underlying resources, but the key is. Keeping that those resources as effectively used as possible. And of course that gives you more benefit for for the the cost that you spend on on on providing a higher memory bandwidth with the latest HPM and higher signaling frequencies and communication bandwidths between the chips using the latest variety of of of NV link and so. Could we zoom out and just ask you to kind of taxonomize the whole chip space if you would? I mean, it's a big question, but I think people are familiar with things like Cerebrus, which obviously has this giant chip and has like a ton of memory on chip. We've seen a number of instances where people are burning the transformer architecture directly into the silicon with varying degrees. I think of flexibility still remaining as they pursue that strategy. And I guess I'm, I'm curious as to how you see the kind of menu of big different strategies, the big different bets that people are making and then also how you see the strengths and weaknesses of each.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.