Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
2 Sept 2026 · 23:55 Latent Space The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
“Um whereas in our case you know we're running the world's largest models and in some ways kind of simple because you know one of our chips has you know order 100 times more memory than one of their chips right so you have two orders of magnitude difference in scale kind of for free right and so I think that's definitely one of the things that was a topic of discussion again independent of of cerebrus um just kind of odd that you know you'd launch a brand new product on on performance numbers of such a small model.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 2 Sept 2026 · 23:55
- Publisher
- Latent Space
Transcript context
…y which is already substantially different than than your baseline GPU and then on top of that there's opportunity to integrate even further like pre-fill and decode disagregation for example or other >> for they didn't actually specialize jalapeno for right >> they did not specialize jalapeno for that but they specialized it for throughput right and so they get a tremendous amount of throughput and so just as a computer architect there's like so many different things you can start to do with that right and then we have you know the you know insane uh latency right like I mentioned up to 10,000 TPS in in CS5, you can start to imagine some some really really cool products that we can build together, right? And that that's what really excites me about having them as a partner. And we're also excited about collaborations on the AI tooling front because much of the the benefits that you know that that they're seeing from the AI tooling infrastructure um everybody probably says this now but like we're obviously doing you know a lot with AI but we're also collaborating very closely with OpenAI right to use their tools to help us also continue to push what's possible in our chip design and our software and all that. So both of those together I feel like you know this is a really unbeatable combination. >> I think one thing that people are talking about like you know I'm trying to get to the disagreements of the hot takes now. So uh people focusing on you haven't really mentioned like power and I I do think that something that seems to be a consistent theme is um you know performance per watt rather than than tokens per second. any variation of of this theme or what are people sort of talking about offstage that you know is more contentious? Well, I think there's a there's a few things that, you know, in terms of, you know, some of the uh the more contentious things like I I would say one of the the themes that came out quite a bit in my discussions at the conference was around the the Grock announcement, right? You know, Grock, obviously, not surprising they have a new chip. In general, I think it's it's awesome that SRAMM designs are becoming, you know, more mainstream now. Um, you've >> been here the whole time. We >> we [laughter] we've been talking about it for a long time. And it's amazing to see, you know, the the industry starting to embrace it, right? Um, >> do they feel different post acquisition? I mean, you've been competing with them for >> a while. >> Uh, to first order, no. I think it's really awesome to see that, you know, SRAMM architectures are becoming, you know, more accessible. you know, even the biggest of the big guys here, Nvidia is embracing SRAMM design, acknowledging that, you know, that uh the traditional GPU designs really can't hit the ultra fast, you know, regimes. A lot of what was being discussed uh offstage was I mean the natural obvious questions is like why did they launch on a 30 billion parameter model? Um, and you know, how come when uh Jensen spent so much time at GTC talking about uh attention FFN disagregation, there was no mention of that. And so I think there's, you know, llion parameter model? Um, and you know, how come when uh Jensen spent so much time at GTC talking about uh attention FFN disagregation, there was no mention of that. And so I think there's, you know, I think that's that's pretty telling, right? Um, >> I mean, there's a separate Reuben, you know, strategy. >> Well, so there's there's there's Reuben, but you know, the LPX itself >> that that was supposed to be where it was. >> It was supposed to be Reuben LPX together, right? Um, if if you guys recall, I mean, the Jensen spent like >> GTC >> GTC like half an hour explaining attention runs here and [laughter] and you know and and the movies run here and so on. And I don't know if this is like a a hot take per per se, but you know, it it it's it's very suspicious that they're um that their product that's in full production, they've only shown performance numbers on a non-disaggregated 31 billion parameter model, right? And I think to me what this this shows is that there's definitely some challenges in running on a nonwafer scale SRAMM design because there's just not enough memory in each of the chips, right? And I think that's you know that that's what's happening and and and we're seeing the evidence of that and in many ways I think it's very much validating kind of the design choice that that that we had right uh if you think about it to run a frontier level model of like let's say a few trillion parameters you need thousands and thousands of Grock LPUs just to hold the the weights. >> Yeah. >> Right. And so when you start to think about it that way, it's like, well, is it surprising that the only performance numbers that they're showing are, you know, on 30B, right? Um, >> so they're going to gradient this into this. [laughter] >> Well, I I think I think it's I think it's the other way, right? I think what's what's going to end up happening is they're going to end up focusing on significantly smaller models. >> All right. Um, you know, if you if you have that limitation in your architecture, then I I think that's that's what ends up happening, right? Um whereas in our case you know we're running the world's largest models and in some ways kind of simple because you know one of our chips has you know order 100 times more memory than one of their chips right so you have two orders of magnitude difference in scale kind of for free right and so I think that's definitely one of the things that was a topic of discussion again independent of of cerebrus um just kind of odd that you know you'd launch a brand new product on on performance numbers of such a small model. But when you peel it back a little bit, it kind of makes sense. I mean, I've been living in this space now for a long time. There's a reason why we needed the wafer scale integration to be able to aggregate enough SRAMM to be able to actually make it useful for large models. >> Yeah. And look, the market's large, right? You have a different market than them. And you know that you you you clearly are the longest running incumbent now in this space. >> No, absolutely. I mean the the the market's large there's a lot of different opportunities for you know m. And you know that you you you clearly are the longest running incumbent now in this space. >> No, absolutely. I mean the the the market's large there's a lot of different opportunities for you know different different hardware to to play different roles. In fact, in general, uh, you know, at Cerebrus, we we we believe very very strongly in a heterogeneous disagregated, you know, uh, ecosystem, right? Not just pre-filled decode disagregation, but, you know, I think we're just at the beginning of what's possible in this space. And with the scale of of of these deployments and of these models, right, inference is no longer just like one workload. There's many many kind of subworkloads within it. And you know, you really want to use the right, you know, tool for the for the problem. You really want to use the right hardware for the problem. And so, absolutely, I think there's there's a there's a spot for all the different types of architectures out there. And, you know, we we've we've chosen to to target, you know, the frontier. Um, >> yeah. Frontier ultra fast. >> Exactly. Exactly. >> Um, I we was going to go into some of the other companies that were, you know, top of town this year, >> but it gives me this gives me opportunity. I want to follow up on one thing which is how do you think about the classes of workloads right um to me the ultra fast frontier workload is just uh like very clearly like one of the fastest growing segments of the entire infrance market right like the fact that like I want it and I cannot pay for enough for it um I it might have been a mistake for openi to offer it even right like but it's good for anyway so uh just in your experience right you said you said you're strong believers in a heterogenous um inference solution okay what are the buckets and um how do how do you see it from your talking to your customers? >> Well, so I I think there's probably two different views here. Um you know, one is the the product view and then the other is kind of the the the technical computer architect view, right? From the product view, it's actually just very simple is that bringing more speed opens up significant opportunities, applications, different use cases, different capabilities, more intelligent models, more intelligent agents. And so the further you can push that, the more and more that you can enable. In fact, there's probably all sorts of things that we can't even imagine that you can build, you know, when you're even faster than what we are calling ultra fast today, right? Which is why we keep pushing that. And then the other angle really just comes down to capacity, right? That's so it's like yes, I want the fastest, but then you need enough to actually be able to to to serve your use case, right? And so it really just is that simple, right? And then finding the right architecture and the right mix of architectures, right, to be able to provide that is is really the name of the game. Now, from the kind of computer architect's point of view, in a lot of ways, it's it's even a little bit simpler, right? Like I was um having a conversation with somebody actually at hot chips about this and they were…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.