Evidence receipt / belief
Published · transcript-backedRamin Hasani: belief
4 Jul 2026 The Cognitive Revolution Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models
“We're talking about a model class that would actually sit on top of, let's say, hardware on a laptop, for example. And yeah, so I would say you need to have that, but I mean, if the default is just giving you that efficiency and that, Like it's ready to go, that's like the choice, that's something that I think Nvidia is trying to propose, like in enterprises, when they go and sell, like the NIM project, Nvidia NIM projects, when they go out there and sell these things, it makes a lot of sense, because you already have like a project, you already have a multimodal model that is loaded on top of your, let's say, the PC that you bought, you know, and...”
Source trail
Everything needed to verify it.
- Speaker
- Ramin Hasani
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 4 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…So does that imply a future where we have just a lot of vertical integration and a lot of coupling? The model will come with the hardware that I buy, and it may not be so swappable in the future because the model is heavily optimized for the hardware and vice versa, such that these things are not so modular in the future as they are today. Or why do you want to change the model? It brings you to that choice kind of question. You're given basically a default, now you want to switch this thing by all means, but why do you want to switch it? If this model, the intelligence layer that is in there, it's not fixed, it's kind of an adaptable kind of system, it is a, let's say, self-improving system, like with a call to your data sets and stuff, you can actually have a platform that does, let's say, full fine-tuning of that system, it is enabled. And there is not just one model that you can load into the system. And there's not just one cloud model or one on-device model that you're going to use. You've got to be able to orchestrate between, let's say, many different instantiation of this model to be able to build application. We're talking about a model class that would actually sit on top of, let's say, hardware on a laptop, for example. And yeah, so I would say you need to have that, but I mean, if the default is just giving you that efficiency and that, Like it's ready to go, that's like the choice, that's something that I think Nvidia is trying to propose, like in enterprises, when they go and sell, like the NIM project, Nvidia NIM projects, when they go out there and sell these things, it makes a lot of sense, because you already have like a project, you already have a multimodal model that is loaded on top of your, let's say, the PC that you bought, you know, and... Why should I switch? Because they already have done all sort of optimizations for me, and it is running extremely fast. Why do I need to change it? And the model itself is tunable. I can use their Megatron kind of framework to actually tune the model. Now, if you don't want to do it, and you want to just choose another model to host it. As I said, this has to be given. Your hardware should be already, you have to be optimizing for the entire open source ecosystem and models that are available, but at the same time, it would give you an advantage, like to yourself and to your customers. Okay, maybe in the last few minutes, how about a practical application of this? You have a blog post on local co-work. No cloud, no waiting, tool calling agents on consumer hardware with LFM2 24B, A2B. So that's a, I think people know, that's a mixture of experts, 2 billion active out of 24 billion total parameters. Let's say I want to make that a part of my life. I am interested in how you would coach me on setting this up. Today I have, for reference, And I, by the way, try to make the, sometimes I try to make the transcript of the podcast something that I can feed to my agent. So you can think of this as partly coaching me, partly coaching my agent. So I've got this deep context database that is the last five years of my digital output. This podcast will be recorded, obviously recorded, transcribed. It'll go into this database. So everything you said, everything I said will be in there and searchable. And it's got all my e-mail and Slack messages and everything. Okay, cool. So now I've got Claude. on my desktop that can call tools locally to get data back, but then it sends all the results to the cloud to decide which of those results are actually the right ones to be looking at. And so far, I've been okay with that. The benefits are certainly worth whatever risk I'm taking, I feel. But I would maybe love to run that data through a local model first so that I don't have to send all my data to the cloud every time. and notably what it gets filtered through, I'm probably still going to end up burning through a foundation model. So it's not going to entirely skip the cloud. But also I'd like to save some tokens too, because I'm going to be going to use Fable whenever I get it back for whatever it's most appropriately used for. So how do I get really good performance on my local computer with this model? Do I need to be doing fine tuning? Do I need a distillation strategy? How do I actually take the base model that you have and recover as much of, let's say, Opus or even Fable performance in terms of searching through, understanding my data as I possibly can? And how much is that? What should I expect in terms of how far I'll be able to push that process? How close to parity with frontier models can I get? So Tell me everything I need to know, and I'll go have my agent do it.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.