High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ramin Hasani: belief

4 Jul 2026 The Cognitive Revolution Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models

“Like you've seen like diffusion is also like a, I would say like this format of a prior that you put on a certain architecture, but it's still like you, there's debate between like diffusion and learning algorithms or diffusion is actually part of the architecture.”

— Ramin Hasani

Source trail

Everything needed to verify it.

Speaker
Ramin Hasani
Attribution
Verified speaker
Claim type
belief
Recorded
4 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, very interesting. Okay, let's look at the other end of the spectrum then. As you work with customers, what are some interesting examples of when Given resource constraints, given the narrowness of the domain of interest, other kinds of bias are actually winning in the architecture search process. Great question, so... For example, if you go to, let's say, biology, and you want to model sequential data in biology, you're talking about DNA data in biology. DNA data from a vocabulary kind of perspective, they're very limited, right? They're not like language that is like maximum kind of amount of vocabulary and stuff. In DNA language, they're very simplified, but... The lengths of those sequences you have to actually process are, let's say, for a human being, or maybe for a bacteria, I heard, it's somewhere between 1 to 100 billion sequences, like basically elements in a sequence. For that long context kind of things that you want to perform, when you do not have that much of a large vocabulary, you don't need attention. So you can actually run on this kind of data. You can run... pure convolutions, pure SSMs, your pure liquid neural networks, like system recurrent neural networks and parallelized version of these recurrences and linear attention for extremely long context, you cannot do it any other way. The reason behind it is because context has become so large that that quadratic cost of attention just kicks in. So let's say on biological data, you would want to have some sort of a structure there. Then there are places like video modeling and stuff, in those kind of places, let's say you might want to have different architectures and even different learning algorithms. You have seen like the success of diffusion, for example, there, right? Like you've seen like diffusion is also like a, I would say like this format of a prior that you put on a certain architecture, but it's still like you, there's debate between like diffusion and learning algorithms or diffusion is actually part of the architecture. You can actually make that connection and you can make that distinction as well. So you have seen on video and seen understanding probably in elements of diffusion would be needed, like we some people still believe that with autoregressive kind of modeling you could actually get there anyways, but we will see, we'll see if that stays true. Then when you're talking about audio signal as well, like if you audio alone, for example, if you're talking about audio alone and language is not part of the whole thing, it's just pure like voices, like voice to voice, let's say noise to clear kind of signal, you know, signal to signal, these are places where recurrent neural networks are still very, very powerful. They're extremely powerful, especially. in the low data regime. I would say models that have a lot more biases in them, like in this smaller kind of regime, when you do not have that much data, this is kind of the place where you can bring a lot of value. And recurrent neural networks and biases in architecture can help you in the low data regime to really fill out that gap with the feedback mechanisms that they have. here you can bring a lot of value. And recurrent neural networks and biases in architecture can help you in the low data regime to really fill out that gap with the feedback mechanisms that they have. The more complex architectures, the more you would be able to handle like the closer the architecture is to the dynamics of the data sets that you're trying to solve, the better of a learning system you're building. In fact, in the physics modeling, like when you're talking about, there was like a class of models like physics-informed neural networks, these are kind of, again, another architecture that we're talking about here. They would be really good for physical kind of simulations. And yeah, so this is the other side of the spectrum that I'm talking about, like the property of the data, the lengths of the context lengths, and all the other considerations that I said would change the architecture by a lot. And then again, you want to scale this to the largest kind of regime, make it, again, unbiased. Again, a transformer would be able to add scale to beat this. But at a smaller scale, transformers would not be able to beat any of the other formats of technical systems that we described.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence