High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“I mean, it's a big question, but I think people are familiar with things like Cerebrus, which obviously has this giant chip and has like a ton of memory on chip.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…e the bandwidth rather than with through NB Link and advanced HBM integration and software optimizations like Tensor MTBLM. So is it, are they actually trying to brute force their way into this? Yeah. I mean at this you really of course want to continue to get improved improvements, peak improvements on HPM bandwidth and chip to chip communication by using the latest technology. But then the key is how effectively do you use that bandwidth, how effectively do you use that communication and do make sure that you don't waste it, right. And so whereas GPU's are often running at maybe 10 to 20% of the the capabilities of the resources, right, the bandwidth and, and the the memory bandwidth and the communication resources, our goal in a Salmonova system is to push that to be 70 to 80% of the peak. And so the idea is, yeah, everybody wants more capabilities from the underlying resources, but the key is. Keeping that those resources as effectively used as possible. And of course that gives you more benefit for for the the cost that you spend on on on providing a higher memory bandwidth with the latest HPM and higher signaling frequencies and communication bandwidths between the chips using the latest variety of of of NV link and so. Could we zoom out and just ask you to kind of taxonomize the whole chip space if you would? I mean, it's a big question, but I think people are familiar with things like Cerebrus, which obviously has this giant chip and has like a ton of memory on chip. We've seen a number of instances where people are burning the transformer architecture directly into the silicon with varying degrees. I think of flexibility still remaining as they pursue that strategy. And I guess I'm, I'm curious as to how you see the kind of menu of big different strategies, the big different bets that people are making and then also how you see the strengths and weaknesses of each. Yeah, that's a really interesting question. I think maybe you can think about it along three different axis, right. So 1 axis is your flexibility versus specialization axis, right. You know, extreme flexibility might be something like ACPU or to some extent maybe a GPU which is this instruction driven execution engine, right. And so can be pretty flexible. But of course you always, always pay overhead for executing instructions, right? And you pay overhead in terms of silicon area and in terms of time, right. And then, you know, on the extreme of that access would be something that would be very specialized for a very specific algorithm, right? And so if that algorithm changed in any way, then that piece of silicon would no longer be useful, right? And, you know, fixing your architecture to Transformers and burning your weights into, into the design would be, might be an extreme case of that, right? But I've learned never to bet against the innovation capabilities of software people and people. And so, you know, I've seen, even though the time that I've been looking at ML and AI, that there's been this tremendous change in algorithms. And of course, now we're kind of fixated on Transformers. But Transformers even aren't just one thing, right? You've got various types of different Transformers. You mentioned state space techniques, you mentioned the fact that the people are coming up with different ways of doing attention, right? And so, so I would be very wary of a sort of kind of fixing any particular algorithm into architecture because then you can't innovate, right? So that's one actor. And so the idea is how can you be completely flexible, but with very, very low overhead, almost no overhead, right? And so the problem with GPU's is they do use HPM so they can run large models, but they synchronize the data movement and the movement of communication of the data between chips, all in software, right? And that adds overhead. And it means that in particular that they have a lot of trouble overlapping computation and communication. And that is in fact the key, right? So what you want to do is you want to communicate, but you don't want to communicate by waiting until you have need to communicate and then you have to run instructions to to move the data. What you want is to construct a pipeline in which the communication is just one component of the pipeline. And so the way to think about this data flow execution is that communication is happening all the time in Yeah, And it's just one of the pipeline stages. And communication is happening for the last piece of computation piece of the computation of the model, while the computation for this piece of of the model is happening in some other stage in the pipeline, right. So it's a classic idea from computer architecture, pipelining and the use of a memory hierarchy to move the data when you need it to where you need it at the right time.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence