Evidence receipt / evaluation
Published · transcript-backedRamin Hasani: evaluation
4 Jul 2026 The Cognitive Revolution Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models
“ficient format of architecture. And it became the double-gated convolution that actually came out of this massive search space, like AFMD, you know, like this original kind of system that we designed.”
Source trail
Everything needed to verify it.
- Speaker
- Ramin Hasani
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…ted inside this linear input varying, or let's say in general, liquid input varying kind of operators, because input dependence is something that we have been talking about this for the last 10 years. That's extremely important. And we see that right now in transformers is also extremely important to have. That input dependence is actually coming naturally with the attention. architecture as well, but in a very, very unstructured way. It's not as structured as, let's say, like what is done in RNNs, in liquid neural networks, in, like in SSNs, in all of those things, you don't see these gated variables like to be very structured in transformers. Now, we actually have Various variants of a convolution operators, various variants of a recurrent operator, and then the attention themselves, there's like there's group query attention, there was the original transform, there's like many different things. Then, on all some of these variations of dynamical systems, and also various types of convolution, I said convolution before, so, but then what we brought, we brought also the liquid computational blocks, like, for example, the gated, the double gated conference, so gating what like a certain type of biases on to like different dynamical systems or different operators. We added them, and probably the space becomes like something around 50 to 100 different operators. And then you want to build hybrid models so that you can reduce. What's the goal here? The goal here is that maximizing efficiency of computation without loss of accuracy. That's kind of the goal that we set. The objective function of the search space is also this. Let's run. You know, let's have like all of our, let's say all of our compute actually thrown at this problem and let's try to see like. how the system is going to design this. And we started doing scaling laws on this. Like we went as a small as a neural network as let's say 10 million parameters to running the scaling laws up to like 72 billion parameter models, you know, like in these hybrid kind of structures. So we have done a massive, early on, early days of Liquid, like 2023 and 2024 has been always like proving out on a certain type of processor, what's the most efficient type of architecture that can come out. And turned out, as all of our, when we put all of our biases away, all the gating stuff that we are putting around operators, they got simplified. So, in Mamba style architectures, you have a bunch of gating mechanisms in, let's say, gated delta nets, you know, gated delta nets, you know, you got, you have like a bunch of operators in there, you know, in linear attention, there are like gated variants. Everybody's like tweaking a certain parameter in the network, like by hand, because in their own experiments, actually, they observe something. Turns out all of this has to go away if you want to get to the most efficient format of architecture. And it became the double-gated convolution that actually came out of this massive search space, like AFMD, you know, like this original kind of system that we designed. ficient format of architecture. And it became the double-gated convolution that actually came out of this massive search space, like AFMD, you know, like this original kind of system that we designed. And this was one of the candidate architectures that came out that turns out to be very good on general-purpose computer that I call CPUs. On CPUs, because CPUs don't have, like they have a special kind of structure, right? On CPU, on all the CPUs coming out of AMD, Qualcomm, let's say Intel, you know, everyone that is building a processor, ARM, all the ARM processors. We try to get to the place where test the different operators, what's that generic kind of structure and let's say computational graph that gives you to the most simplified, no, you know, no added hand tuned, like there's no hand tuned features to anywhere, you know, it's just literally coming out of that, the test, the semantic test that we've done. And this became like the de facto kind of architecture LFM to structure that we actually announced. But while exploring these things, like we've observed like many different candidates also popping up. like there's like too many things. And then we try to really like figure out, for example, for let's say an NPU and neural processing units that is like inside an AI PC powered by AMD or powered by Qualcomm or powered by Intel, which of these variants would be like a better neural architecture that gives those hardware providers and silicon portals a boost on the amount of efficiencies that they are unlucky, or let's say a speed of competition, latency that they get, memory footprint that they get, while having a computational graph that doesn't sacrifice any format of quality. So this was like the whole thing. Removing all the human bias with a systematic approach or two architectures, even our own biases, you know, and the only thing that actually remained in elephant 2 on CPU competition is this double-gated As I told you, we have nested computation in liquid neural networks originally. This nested format of computation was something that actually became very interesting to be part of these, a very simplified kind of neural networks that we built. And the fact that we have unstructured 1D convolutions as the layers of choice, 70 to 80% of our networks are structured by these gated convolutions that we have, double gated convolutions that we have. They're extremely simplified and they replace attention. They reduce the computational complexity by a lot. They reduce the memory footprint by a lot. They maximize kind of the speed of computation. by a lot and at scale as well. And at the same time, on the quality, I mean, you have seen some of these models are actually extremely kind of competitive to the transformer-based alternatives. I think you asked a bunch of other follow-up questions as well, but I would pause here for any follow-ups here. Yeah, let's get into use cases in a minute. I'm definitely interested in that. And I also want to... talk about the future of hardware and what your work implies for the future of hardware. But I think it would probably be helpful for a lot of people, including myself. Although this is something I've studied in some depth, I still would like to crack it better than I do. This concept of gating, I'll give you my rough and ready understanding and then you can improve it and deepen it. I especially notice this with Mamba. when I went down that rabbit hole and became very excited about it. seems that the kind of trick that is played over and over again with these gating mechanisms is we want the transformation that is done on the data to be input dependent. So it's not enough to learn a transformation. We want to learn a transformation, but then also have some relatively, typically the gate is relatively low dimensional. Sometimes it could just be like a scalar that's applied to that transformation. It could be obviously more complicated than that. But it's some relatively simple mechanism that says, for this learned transformation, here's how we're going to modify it given the input currently under consideration. And that Seems to be great. It seems to unlock a tremendous amount. Help me understand more anything you think I'm missing there, or help me deepen my intuition for why that is such a powerful and recurring theme in all these different architectures.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.