Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“It turns out for the kinds of capacitors we use, you see variations that are on the order of, you know, sort of 10 parts per million, right. So giving you levels of precision that are in the neighborhood of 20 bits of precision, which it turns out is well beyond what we need for the quantization kinds of, you know, levels that we care about, which are typically at the level of eight bits and you know, higher than that in some cases.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…are tolerant and that they are, right, We're able to do things like quantization, as you pointed out, but really we're doing those kinds of things with very carefully represented noise sources, right? Quantization noise is a very carefully represented noise source and one which you need to be able to properly represent throughout your layers of abstraction and digital. You know, quantization is something that we do have ways of building robust abstractions for, but this analogue noise is one that that we really don't. And that's why essentially what you do, you know, when you do digital compute is you say, hey, listen, there's all sorts of noise that analog might lead to. But we drown all of those out by thinking about, you know, the signals being a 0 or A1. So that that's the dominant source of noise. And that's all that I now have to represent over my layers of abstraction. Everything else basically doesn't matter. But because this is a very well represented form of noise, quantization noise, I know how to deal with it at the algorithmic level. And I can apply all my algorithmic techniques. And that's what the industry has done very successfully. But now as you want to sort of leverage analogue, you still need those levels of abstraction. That's the key to achieving systems at scale and and you know, systems that you can build architectural abstractions and the software abstractions on top of. So I would say that the need to be accurate and precise is still, you know, kind of brutally high. And so that's that's very, very important. Now you also ask the question of OK, then how precise is your is your approach? Because now you know, I'm telling you we need to know that and understand that very well. And it turns out that the dominant source of noise that we have in our approach based on these capacitors, there could be many sources. There's the electronic, you know, the discrete nature of electronic charge causes noise and things like that. But it actually turns out the dominant sources, the variability of the capacitors that we we can fabricate on a chip. Now, it turns out that those capacitors are really critically dependent on geometric properties. So basically the distance between 2 metal wires. And it turns out that geometry is really the one thing we can control very well in CMOS processes. It turns out we use this processing approach called lithography that gives us very precise geometric control. That's the reason we can build, you know, 532 nanometer transistors. Turns out we don't need anywhere near that precision for the capacitors that we use. But, but it's really because of this alignment with this geometric control that this particular approach has that allows it to be brutally accurate in the ways that you need it to be through all of these layers of abstraction to be able to scale up. We've measured these things in a lot of detail. It turns out for the kinds of capacitors we use, you see variations that are on the order of, you know, sort of 10 parts per million, right. o scale up. We've measured these things in a lot of detail. It turns out for the kinds of capacitors we use, you see variations that are on the order of, you know, sort of 10 parts per million, right. So giving you levels of precision that are in the neighborhood of 20 bits of precision, which it turns out is well beyond what we need for the quantization kinds of, you know, levels that we care about, which are typically at the level of eight bits and you know, higher than that in some cases. So we've, we've had to characterize these things very carefully because the noise does matter as you're trying to build these abstractions all the way up. And that's the level of precision we've we've gotten here, which is what made, which made this approach so practical and where we've been able to now scale it and demonstrate it across, you know, all sorts of chips and systems. I note that you have spent quite a bit of time getting a neural net onto one of your chips and onto I think a laptop like a edge edge device, right? Trying to get into edge devices which are more sensitive to power consumption and more sensitive to, you know, I you, you just don't have the affordances that you have in a data center, so to speak. So what would be the comparison be like, you know, if you were to use a normal GPU versus A1 of your chips, What is a comparison on like, let's say energy savings or like how do you compare the 2? Yeah, that's a great question. And and I think it really points to the fact that if you want to now leverage this, you know, fundamentally new technology, analogue, it's not enough to just build that technology and make it robust. You end up having to build the entire architecture and the entire software around, you know, harnessing and extracting it's full efficiency. And the reason I say that is I can give you 2 answers. One is at the level of the core technology, the thing that this analogue computing engine does, what level of efficiencies do we have? fficiency. And the reason I say that is I can give you 2 answers. One is at the level of the core technology, the thing that this analogue computing engine does, what level of efficiencies do we have? Turns out, as you guys know, the bulk of the operations that we do in, in, in AI compute our matrix multiplications or matrix operations, tensor operations. So essentially what this engine does is matrix multiplies. And at that level, I think we've now publicly disclosed and you know, we've got silicon and you can come to our labs and many of our customers and partners have and they see it. You know, we're, we're basically doing an 8 bit compute at 150 tops per Watt in a 16 nanometer technology. So just as a point of reference, the best digital matrix multiplies will give you sort of like 5 pops per Watt in that technology node. So this is 30X better at the level of that that core technology. One of the very exciting things for us is that we've taken our technology from 16 nanometer CMOS and and been able to scale it to very advanced nodes right now. And, and as we've predicted and as we've seen in our from our previous chips, the energy efficiency advantages just just scale. And the reason is because it all depends, as I mentioned, all this geometry. And so as we move to Finder and Finder CMOS node, your geometric control and densities are getting better. And, and so we, we benefit from that in, in our analogue approach as much as we do in digital approaches. But the important point I wanted to make is that, OK, that's just the efficiency of running this core matrix operation. But, you know, there's all of this other stuff that happens around this. There's operators that are not matrix multiplies, there's not linear operators, activation function, softmax, on and on and on, right? And then there's all the infrastructure you need to actually run this in a programmable way, right? Some models are big and some are small and some are convolutional and some are Transformers and some are, you know, have layers that that need to decide how to route, you know, sort of tokens to one expert or another. So also to compute that needs to be integrated that you need to make programmable. So now the problem is you've taken this, you know, core operation and made it 30X lower energy, basically made it's energy almost zero. Everything else now needs to also be addressed, and that includes the architecture, the entire memory system, the way that the software executes on it. And I think what's really been exciting for us is that big breakthrough actually happened in the lab, the switch capacitor approached in memory computing back in 2017, right? Our lives since then have really been about how do you build architecture and software and integrate these into standard workflows and so on so that you preserve, you know, that efficiency advantage at the level of the full system end to end executions. And so you always incur overheads because of all these other things you have to do.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.