Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“I'm not a bidder, but I didn't have the model way back then. Because I think that, you know, does help us drive really the the fundamental architectural approaches for for programmability and scalability, which will serve us into the future.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…fficiency. And the reason I say that is I can give you 2 answers. One is at the level of the core technology, the thing that this analogue computing engine does, what level of efficiencies do we have? Turns out, as you guys know, the bulk of the operations that we do in, in, in AI compute our matrix multiplications or matrix operations, tensor operations. So essentially what this engine does is matrix multiplies. And at that level, I think we've now publicly disclosed and you know, we've got silicon and you can come to our labs and many of our customers and partners have and they see it. You know, we're, we're basically doing an 8 bit compute at 150 tops per Watt in a 16 nanometer technology. So just as a point of reference, the best digital matrix multiplies will give you sort of like 5 pops per Watt in that technology node. So this is 30X better at the level of that that core technology. One of the very exciting things for us is that we've taken our technology from 16 nanometer CMOS and and been able to scale it to very advanced nodes right now. And, and as we've predicted and as we've seen in our from our previous chips, the energy efficiency advantages just just scale. And the reason is because it all depends, as I mentioned, all this geometry. And so as we move to Finder and Finder CMOS node, your geometric control and densities are getting better. And, and so we, we benefit from that in, in our analogue approach as much as we do in digital approaches. But the important point I wanted to make is that, OK, that's just the efficiency of running this core matrix operation. But, you know, there's all of this other stuff that happens around this. There's operators that are not matrix multiplies, there's not linear operators, activation function, softmax, on and on and on, right? And then there's all the infrastructure you need to actually run this in a programmable way, right? Some models are big and some are small and some are convolutional and some are Transformers and some are, you know, have layers that that need to decide how to route, you know, sort of tokens to one expert or another. So also to compute that needs to be integrated that you need to make programmable. So now the problem is you've taken this, you know, core operation and made it 30X lower energy, basically made it's energy almost zero. Everything else now needs to also be addressed, and that includes the architecture, the entire memory system, the way that the software executes on it. And I think what's really been exciting for us is that big breakthrough actually happened in the lab, the switch capacitor approached in memory computing back in 2017, right? Our lives since then have really been about how do you build architecture and software and integrate these into standard workflows and so on so that you preserve, you know, that efficiency advantage at the level of the full system end to end executions. And so you always incur overheads because of all these other things you have to do. on so that you preserve, you know, that efficiency advantage at the level of the full system end to end executions. And so you always incur overheads because of all these other things you have to do. We want to make sure that we maintain that kind of ratio of overhead so that, you know, a 30X advantage in the fundamental compute still gives you, you know, sort of order of magnitude advantages at the full system level. And that's where really all of the innovations have been since, since that initial, you know, breakthrough in, in 2017. One question I have, if you had access to GPT 5.5 in 2017, would it have accelerated your work because you're the fundamental breakthroughs already came then and you've been building out the harness and you know, all of the supporting infrastructure. Would it have accelerated your work if you had access to one of these models in 2017? You know, it's a great question, right? I mean, one way that way I can interpret your question is to say, hey, listen, you know, the models are always moving. If you do the model and where it will be 5 years from now, maybe you could have just built that architecture for that model, you know, immediately rather than going through, you know, sort of the support that you need for all of the models that came, came in between. You know, the interesting question here, Prakash, is even if I had GPT, you know, the next version, right? You know, there's another version coming after that. And so the, the, the architecture does need to be built from the ground up in a way that that supports algorithmic innovations, architectural model architectural innovations. And so I think that, you know, the, the work that that's gone on, even as we've tried to onboard, you know, models in that interim and, and make them run efficiently is all very productive work. It helps to drive a general concept of how do you build very programmable and scalable hardware using now these new analogue based techniques for the fundamental technology. So that, you know, that's kind of the way I see it. I'm not a bidder, but I didn't have the model way back then. Because I think that, you know, does help us drive really the the fundamental architectural approaches for for programmability and scalability, which will serve us into the future. Where do you think is the first device that we'll see like, you know, in that a consumer might see with your technology in it? Yeah. ogrammability and scalability, which will serve us into the future. Where do you think is the first device that we'll see like, you know, in that a consumer might see with your technology in it? Yeah. So the, the first devices are going to be, you know, client computing devices, that's VI powered laptops, you know, also desktops and, and workstations and, and things like that. And the reason is, you know, as we started to really build out this, you know, technology into a real product, hardware and software and all of those sorts of things. Back when we started the company that was back in 22, you know, the place you really needed energy efficiency was, was at the edge. And this was right around the time chat 2 PT had just come out where we were deploying models in the data centre and, and seeing all sorts of challenges related to cost related to privacy security. So there's a big industry push to try to move these models to the next adjacent device by which we access the data center, these client platforms. And so that's where we found a lot of partnerships and a lot of industry demand and interest and and that's where you'll see the first products. Now what's happened in the meantime is, you know, back in 22, I'm not sure that energy efficiency and that we spoke about it was the critical thing in the data center, but boy is it now. And so one of the things that end charge has been and doing very carefully and, and thoughtfully is, is working together with the right partners to now bring that level of energy efficiency to really solve the hard constraints that we face in terms of power efficiency in the in the data centre. And that requires different kinds of architectures, but ones where we're clearly seeing this fundamental technology and, and the efficiencies it brings can be designed to to really address in a transformative way. Do we have a, before I go order a Mac mini, do we have a timeline or, or a Mac studio for that matter? Do we have a timeline for when something like this becomes available? And you know, do we have a price point? Do we have a sort of predicted tokens per second at a given model size? Like this may be a little early, but I want to, you know, kind of do my side this side by side against the Mac studio. That might be my other default path. So the, the, the chips and their availability to you to be able to use them in applications is something that's going to happen together with our partners and and the laptop platforms that they'll deliver to the market. So I'm not going to speak to their timelines and so on because of the various ways that they think about marketing these products and and strategies that they have around that, but which we've been very active and engaged with them on. But to give to answer some of your questions right like our our first products for that client computing space are. Processors that provide 200 tops of AI capability. I mean that's the kind of capability that, you know, sort of just a couple of years ago you would have had and you know sort of or even today, right, you really haven't and sort of 150 Watt GP us that kind of thing.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.