High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Ramin Hasani: evaluation

4 Jul 2026 The Cognitive Revolution Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models

“You don't even need some of those hybrids. But hybrid is just boost that accuracy because as we see, as I said, the ON2, like basically the complexity, that computational complexity at a certain level, at a certain scale, it is needed for us to really get to those performances that you want to.”

— Ramin Hasani

Source trail

Everything needed to verify it.

Speaker
Ramin Hasani
Attribution
Verified speaker
Claim type
evaluation
Recorded
4 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah. You remind me a little bit of Ali Behrouz's, the illusion of architectures. If we have time, maybe we can touch on nested learning. Okay, well, let's stay focused on your work. for the moment at least. So in LFM2, this is the result of this architecture search. And the kind of surprisingly simple thing that comes back is some reduced but still critical number of attention layers. And then the other layers are, and this is, we've seen this kind of with SSM attention hybrids as well, but I think the kind of surprising revelation from the result of this search process is the non-attention layers can actually be extremely simple as long as they have some gating. So you've got the gate and then you've got just a real simple convolution that I think I read only considers a very short span of tokens, right? Is it like just the short comes back? Yeah. And So if we were going to update the headline, attention is all you need from however many years ago now, we would maybe say attention is something that you really do still need, at least at a certain scale, but you also need gating on your other layers. But you don't actually need anything like super crazy, fancy, sophisticated, the state space model, all that kind of stuff. It turns out that keeping the gate is actually the part that really drives the most value. And then you can have a really simple mechanism behind the gate. And then you can have 70% of that and 30% of attention and subject to obviously resource constraints that ended up being the winning formula. Am I getting anything wrong there? No, I mean, you're touching on the right things. And this is basically, it's like the regime you're operating and the goal of your system. Are you trying to build super, are you trying to build like the most powerful version of the AI system? You need the most unbiased version of an algorithm. Now, attention is an extremely rich, unbiased format of algorithms. Even if the computational complexity of attention is N to the power two, maybe we really need N to the power two to really get to that kind of level, you know? And maybe And we need more complex architectures. We've always tried to actually reduce the complexity of architectures for the sheer purpose of the fact that we are resource constrained. As humanity as a whole, we are resource constrained right now. So I would say the discoveries that we have right now, it just shows that there is a gradient on architecture that you can follow as you scale models. The gradient that you're following is the fact that for smaller kind of models and specialized models, you can put as many biases, like these scaling mechanisms that you're bringing in. And you can play around with as many operators of interest in your computational graph. And it is going to work and it is going to give you some sort of a boost if you're really maximizing for like linear let's say linear time complexity, like you want to implement linear attention systems, you know, like just the fastest kind of, if the speed is like so important and you're actually wanting to even sacrifice a little bit quality, you can bring in like linearity and the whole system could be linear, you know? You don't even need some of those hybrids. But hybrid is just boost that accuracy because as we see, as I said, the ON2, like basically the complexity, that computational complexity at a certain level, at a certain scale, it is needed for us to really get to those performances that you want to. The larger the network becomes, the more unstructured you can make it. That's kind of the learning from that whole kind of algorithmic approach that we started designing in all architectures. Yeah, very interesting. Okay, let's look at the other end of the spectrum then. As you work with customers, what are some interesting examples of when Given resource constraints, given the narrowness of the domain of interest, other kinds of bias are actually winning in the architecture search process.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence