Evidence receipt / observation
Published · transcript-backedSpeaker unverified: observation
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“What you want is to construct a pipeline in which the communication is just one component of the pipeline. And so the way to think about this data flow execution is that communication is happening all the time in Yeah, And it's just one of the pipeline stages.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- observation
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Could we zoom out and just ask you to kind of taxonomize the whole chip space if you would? I mean, it's a big question, but I think people are familiar with things like Cerebrus, which obviously has this giant chip and has like a ton of memory on chip. We've seen a number of instances where people are burning the transformer architecture directly into the silicon with varying degrees. I think of flexibility still remaining as they pursue that strategy. And I guess I'm, I'm curious as to how you see the kind of menu of big different strategies, the big different bets that people are making and then also how you see the strengths and weaknesses of each. Yeah, that's a really interesting question. I think maybe you can think about it along three different axis, right. So 1 axis is your flexibility versus specialization axis, right. You know, extreme flexibility might be something like ACPU or to some extent maybe a GPU which is this instruction driven execution engine, right. And so can be pretty flexible. But of course you always, always pay overhead for executing instructions, right? And you pay overhead in terms of silicon area and in terms of time, right. And then, you know, on the extreme of that access would be something that would be very specialized for a very specific algorithm, right? And so if that algorithm changed in any way, then that piece of silicon would no longer be useful, right? And, you know, fixing your architecture to Transformers and burning your weights into, into the design would be, might be an extreme case of that, right? But I've learned never to bet against the innovation capabilities of software people and people. And so, you know, I've seen, even though the time that I've been looking at ML and AI, that there's been this tremendous change in algorithms. And of course, now we're kind of fixated on Transformers. But Transformers even aren't just one thing, right? You've got various types of different Transformers. You mentioned state space techniques, you mentioned the fact that the people are coming up with different ways of doing attention, right? And so, so I would be very wary of a sort of kind of fixing any particular algorithm into architecture because then you can't innovate, right? So that's one actor. And so the idea is how can you be completely flexible, but with very, very low overhead, almost no overhead, right? And so the problem with GPU's is they do use HPM so they can run large models, but they synchronize the data movement and the movement of communication of the data between chips, all in software, right? And that adds overhead. And it means that in particular that they have a lot of trouble overlapping computation and communication. And that is in fact the key, right? So what you want to do is you want to communicate, but you don't want to communicate by waiting until you have need to communicate and then you have to run instructions to to move the data. What you want is to construct a pipeline in which the communication is just one component of the pipeline. And so the way to think about this data flow execution is that communication is happening all the time in Yeah, And it's just one of the pipeline stages. And communication is happening for the last piece of computation piece of the computation of the model, while the computation for this piece of of the model is happening in some other stage in the pipeline, right. So it's a classic idea from computer architecture, pipelining and the use of a memory hierarchy to move the data when you need it to where you need it at the right time. stage in the pipeline, right. So it's a classic idea from computer architecture, pipelining and the use of a memory hierarchy to move the data when you need it to where you need it at the right time. And so the nice thing about these AI models is that you do have a graph of computation. And the whole idea of data flow is to take that graph of computation and map it onto the machine in a spatial way, such that you keep all the pieces of the model operating at the same time on different components of the computation that needs to be done.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.