High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

4 Jul 2026 The Cognitive Revolution Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models

“Now, excuse me, what I What sort of jumped out at me in terms of your approach is that you've developed a architecture search process where the promise to customers is not that, hey, we developed this one paradigm and the old calculus teacher used to say, when all you have is a hammer, everything looks like a nail.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
4 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…So continual learning would be a system that is continuously receiving new data and re-tunes itself. So it changes also the parameters of the system. Liquid neural networks are, the dynamics of the systems are input dependent. So that means new data comes in, the dynamics of the system would behave to those inputs, but the parameters of the systems, like any other neural networks, is fixed. It's just because the format of every neuron or every node in a liquid neural networks was like a differential equation, you would have a different dynamics. So dynamics, it's another dimension that you add to a number of parameters. It allows you to compress more information, compress kind of more knowledge. That's why with the smaller instance of the models, there's no free launch or there is no magic here like you're doing like also with liquid neural networks. It's just that the axis of dynamics has been something that we added to the neural networks for what? For like being more adaptable. So I'll give you a very tangible example here. Imagine you're driving and all of a sudden it starts raining. So depending on the format of the rain, your autonomous driving system that is actually taking control of the car could react to those kind of driving scenarios. If you have a traditional neural network that hasn't seen that kind of environment, it might actually get biased, because that rain that actually hits the camera, it's noise, a certain type of noise, or a certain type of adaptation of the input that is coming. But the reality hasn't changed. The confounding variables of the whole environment hasn't changed, right? So that's why liquid neural networks, because they react to the input completely differently, they absorb the input, they apply the low-pass filtering on top of the input, they are much more adaptable to those kind of scenarios. That doesn't mean that they change the parameters of the system to be more adaptable, because there are two axes here. One changing the parameters of the system, the other one is changing dynamics of the system. I would consider continual learning to be attributed to where you continuously also change the parameters of the system. In liquid neural networks, we don't do that. Gotcha. Okay, helpful. Okay, so let's fast forward to the present. You guys are now in the market working with customers. One notable conversation I heard with one of your customers was with the CTO of Shopify on the Latent Space podcast, who had some very good things to say about you and your technology. I was struck by the fact that you've gone neutral in a sense. Now, excuse me, what I What sort of jumped out at me in terms of your approach is that you've developed a architecture search process where the promise to customers is not that, hey, we developed this one paradigm and the old calculus teacher used to say, when all you have is a hammer, everything looks like a nail. So you're explicitly promising to people that even though we came from this lineage of this particular kind of network, we're not just going to blindly apply it to your problem. Instead, we've created this sort of higher abstraction or more meta process for searching through architecture space to find the thing that's going to work best for you. And the two notable details of that are, one, proxy metrics you found don't work that well, just measuring perplexity or whatever, you found you needed to actually go farther and test models on the actual downstream tasks that they're going to be asked to perform. And second, hardware in the loop, actual target hardware in the loop to test the architecture subject to the very real and physical constraints of the robot, the sensor, the phone, whatever it's going to be running on. Maybe I understand that the LFM model came out of that, but maybe before we even get to the LFM, could you sketch out the range of different types of problems that we're putting into this architecture search? And then maybe also a subject that implies, of course, what kind of hardware are we targeting? And then what kind of different architectures are winning for different kinds of problems under different kinds of constraints? Absolutely. So there's a system that we developed in-house, we call it automated foundation model design, you know, AFMD, you know, like that's a meta-learning system. that puts the hardware in the loop and then tries out many different operators with an evolution strategy. The criteria is an evolution strategy, optimizing for a couple of things. Optimizing for memory consumption on that device, optimizing for latency, optimizing for speed, while no sacrifice on quality. When we talk about quality, perplexity is not the measure. It's actually the application, the downstream applications that we care about. It's not just also public benchmarks. We're talking about 100 different benchmarks. So the problem space becomes, from a meta-learning perspective, becomes like a very, very complex kind of problem. Now, I'll tell you that, why did we take this approach, like to design an architecture? Why? Because we wanted to remove all the human biases early on as we are actually building architectures. One of the things that we realized culturally at companies, this is what I can tell you, even at largest foundation model labs in the US right now, Entropic and OpenAI, there are a bunch of people, people that are coming from the science, I call them the Avengers of the architectures, or Avengers of the post-training, or let's say pre-training. These groups of people, there is usually like a very small set of people that are calling the shots on like, oh, you know what? You're going to tweak this portion of this architecture so that it performs better. Why? Because in my personal experiences, it has started working better. If you're really, truly think, and this is something that is broken in all the foundation model labs. You cannot say that somebody has a fix to this, but now the recursive self-improvement kind of process is actually fixing for that, because now people are just finally realizing, you got to give it to the algorithms. You have to be better lessons. You got to be giving it to a systematic way to actually find out what is the true architecture for the problems that you want to solve. You can build like a general-purpose computer. The insights that I shared with you in the format of the scaling laws of neural networks is coming out of our massive exploration of the space of architectures. You know, the fact that in the smaller kind of category of models, some biases on the architecture help. In the larger instances of the models, you don't need to bias the systems. You can actually go pure convolutions, pure transformers, and you can do pure, let's say, reconnaissance that are like very simplified. You don't need to add any form of kind of, specialized kind of treatment like gating there, gated delta nets, and then there's the Mambas and there's the Jambas and there's like, there's a lot of different variations of these architectures that are coming out. What you want to do though, you want to be completely unbiased.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence