High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ali Behrouz: belief

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“Like different aspects to answer this question from the technical point of view, I I think generally the entire literature, most part of the literature in the, you know, past 40 years is built on a paradigm that says we have a pre training or generally training phase and we have a test phase.”

— Ali Behrouz

Source trail

Everything needed to verify it.

Speaker
Ali Behrouz
Attribution
Verified speaker
Claim type
belief
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…uture AI would be like that's different than that? And like how you know, does it look more like another person but with AI advantages? Or does it, does it look like an LLM with like weaknesses patch? But I don't know, like what what is your kind of when you dream of a 20-30 AI that you're, you know, working closely with on a on a daily basis? What do you envision? Like different aspects to answer this question from the technical point of view, I I think generally the entire literature, most part of the literature in the, you know, past 40 years is built on a paradigm that says we have a pre training or generally training phase and we have a test phase. But the point is if we have a continual learning, we can see that, you know, there are like a lot of recent studies about continual learning, how we can do that and how we can like overcome about a lot of challenges. But the point is a true continual learner doesn't have a test and train time. So potentially if we hear that name in any like design choices, potentially it means that it's not a true continual learner because there is no test, there is no train time. And so the question is, is it like a uniform process for the model or not? And my personal opinion is that we still need at least 2 phases. And how it works is that we should have one phase that the model is active so it actively receive information. It's whether through the user query, for example, or for example, it might be about vision models, word models, or anything similar. But the point is the model receives some information and generally performs some computation on the input data and it's active at that point. But on the other hand, there is another phase that the model does not have to wait for the inputs data. It might not receive any input data, it's completely blocked from the word outside of it. But the question is, even at that time, should the model be static without like any performing computation or doing something? Or the model needs to start thinking about some process, thinking about the data that it has inside its parameters and so on and so forth. So I think we can break the process. As I mentioned in two parts, 1 is the active phase and another one is another phase. Potentially we can call it like a sleep time because there is no inputs, but still the, you know, artificial brain or generally like that, that the model itself is trying to perform some computation. And so I think that's a good way of of defining different phases in in this direction of continual learning. And then I think good model is a model that performs very well in both sides. It should receive the information properly, encode it, process it and understand it in the best way possible. And on the other hand, when it goes to the sleep time, it should also like start processing what it has learned before and use that for self improvement. And so that's, I think an ideal model should do from the, you know, technical point of view. But on the other hand, I think there are a lot of challenges. The models that we know right now are, are very large. So even a simple, for example, academic papers that is like presenting a new LLM or architecture or something like that needs to perform some experiments on models with like billions of parameters, 1 billion, two billion or something like that. rs that is like presenting a new LLM or architecture or something like that needs to perform some experiments on models with like billions of parameters, 1 billion, two billion or something like that. And generally, that's a very large model, requires a lot of like computation if you want to keep that model, you know, updated over time. And it needs some techniques to somehow make this process possible. And generally for us, when we were thinking about like this direction, it was a time that some ideas about nested learning started. Because generally if you think about like nested learning, we can see that at each timestamp that we have, we don't have to update everything. We just need to update just a small subset of all parameters. And so that potentially if they to overcome the challenges about the efficiency.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence