High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute

“I have AI, have my own test. And it's been, you know, every time we try it, it fails and I'm like, OK, another, another one doesn't work right.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
26 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…which matters a lot. And it's never quite seemed super compelling to me, especially I think for the, you know, tying back to our very first conversation, why would I want to do it in the first place? And one big reason I'd want to do it would be to search through my own locally available data that all else equal, I would rather keep private and not have to send over the wire. But if it's going to take a super long time to do the prefill to, you know, evaluate the all those records that I have to, to power a search, then it doesn't feel that awesome in the, you know, in the broader context of my stack. So I haven't quite got over the hump where I'm like, this is really going to solve a problem for me. And I'm, I'm still the the question of, yeah, pre time to 1st token and tokens per second are the, the ones that I'm going to be watching most closely for as I wait for a threshold where it feels like now I actually want to do it. 10 billion parameters, right? They, they, they're down to 10 billion parameters at that level. You know, even with the most advanced models today, the models are not like super, super smart. It has to be a model router of some kind. You have to like take the query, make a decision on whether you can handle the query or hand off, make a, you know, make a query to a larger model. So I, I, I kind of want to see where Apple comes out in this because they're the ones best positioned for this kind of like data center plus edge, kind of, you know, handling the query between the two. And you know, they haven't done well so far. Let's, let's, let's see what happens. Maybe maybe the hardware division, but you know, it's a big change, right? They had, they had, they've had the NP us on on device for 4/4 or five years now. We haven't really seen real kind of edge in a computing from Apple yet. I've, I've read through a lot of Apple patents, by the way, they have a very like structured process on hitting a performance window on the device. So they have like they degrade models to fit within the memory constraints, within the latency constraints of the device of the customer. It's a very structured process. And I'm sure Naveen has to, you know, and charge is doing that too. As he says, they're going to have to squeeze the models in and that's the entire harness that you require around the chip, you know, to, to figure out what kind of model is going to work within the latency and performance constraints which the customer expects. So increasingly we are getting a lot of power in not that many billion parameters, right? I mean, the, the Gemma 4 series has certainly pushed that frontier once again. And it's tempting. Every time I see one of these new things, it's tempting and I kind of reread the analysis and I'm like, is it, is it quite there? I haven't quite for the hump yet, but it it might not be too far off. Maybe one more turn of densifying intelligence and you actually could get to a point. I certainly don't need, you know, that that first, as Anna was describing earlier, that first filter of data doesn't have to be super smart. It just has to be somewhat smart to get the, you know, to kind of flash everything, so to speak. t that first, as Anna was describing earlier, that first filter of data doesn't have to be super smart. It just has to be somewhat smart to get the, you know, to kind of flash everything, so to speak. So you know most possibly solving IMOIMO problems on their on their laptops, right? You know most most of it is emails and you know moving data from one place to another if the models on device get good enough and fast enough. You know watching computer use, watching GPT 5.4 or 5.5 do computer use on a computer is very frustrating, right? You, you watch it, make the mistakes. I have AI, you know, everyone has their test. I have AI, have my own test. And it's been, you know, every time we try it, it fails and I'm like, OK, another, another one doesn't work right. So I, I, I found Anna. You know what, what Anna has done is quite interesting. test. And it's been, you know, every time we try it, it fails and I'm like, OK, another, another one doesn't work right. So I, I, I found Anna. You know what, what Anna has done is quite interesting. They they seem to be in the same space as glean as well right now because they're going after enterprise search. And I wonder to what extent you require a large sales team for that, whether this kind of plug in concept works or, you know, in order to implement ceramic at a large firm, you probably need to go in into their firms, you know, VPN, etcetera, and into the inside inside the firewall. And a lot of firms have concerns about having AI, you know, prompt injectable AIS operate within the enterprise firewall. I think I think a lot of a lot of enterprises are still trying to get over that, over that humble security. Nematron Nano 3. Is it going to get prompt injected? How does a prompt injection work? You know, I, I, you know, this morning, one of the opening eye guys, he showed, he put a screenshot of him checking his e-mail using chat DBT 5.5 and it has the number like 4 different prompt injections which are coming into his e-mail box. So, and, you know, obviously they are, they are a huge target for hackers and they've, so he, on a normal morning, you wake up, four different prompt injections are coming in and, and, you know, meanwhile we just tell our AI to read our e-mail, right? That's, that's what we all do. So I wonder to what extent like this issue of prompt injection can be solved in order to enable like businesses, enterprises and people to use these things without, without so much worry, right. That's probably the biggest reason that I use Claude is that I perceive it to be most robust to that kind of stuff. I guess there's also just the general vibe that it seems to be quote UN quote better and hard to define ways. But when I think about like, OK, GBT 5.5, I, I, I need to go check that prompt injection stat before I would put it in the same place that I currently have Claude and I am, you know, I'm attracted to some of its cleaner, arguably more ethical behaviours, but that prompt injection thing does kind of concern me given the level of access that I've given to the agent now. So, yeah, it's crazy to think that they're already getting multiple a day, multiple per day. And, and also not like, not just like, oh, you know, I want to know stuff on this guy's, you know, laptop. It's like extract the environment variables from, you know, his GitHub repos on, on, on the, on his local device. Like scary, scary stuff, right? Like if you had, you know, the environment variables for one of them, like you could do a bunch of stuff on their on their repo, you could extract the model weights probably, right? So scary, scary stuff. So yeah, that does sort of suggest a separation of concerns approach to that. You might imagine when you have Nematron reading your e-mail and filtering to provide relevant context back to some smarter model. Maybe it just. Doesn't have any other tools, you know, you can imagine that kind of, that's basically how, you know, I guess a lot of architectures work right.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence