High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Ali Behrouz: evaluation

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“For example, if if it's a mathematical rule saying that you know, just any mathematical rule that we can have, we start with some specific examples and just memorize them. And then at some point, we just neuralize our understanding of all those examples, remove all those examples in our brain, and replace all of those memories with just one single memory that can describe everything that we have learned so far from that concept.”

— Ali Behrouz

Source trail

Everything needed to verify it.

Speaker
Ali Behrouz
Attribution
Verified speaker
Claim type
evaluation
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…Generally, the main idea as we discussed earlier was that if we have a truly continual learner model, then there is no test and train time. At the other hand, we need to have one active time wherever that the inputs is coming in an online manner and also the time that is the time we don't we don't have any inputs. So the model is not actively receive information from the outside, but it doesn't mean that the model should be static. It means that the model just doesn't get inut, but it can have some internal computation to improve itself. And so that's a very, very general concept. We can incorporate more and more components to the sleep time that we have. So it doesn't have to be just these two specific part, but you know, these two were really relevant to the research that I'm doing. So we just like did that, but potentially it can include any other form of self improvement and so on so forth. So that's just one, one way of breaking the life of a continual learner into active time and sleep time. But what we have in the sleep time right now, which again, as I mentioned, we can incorporate more components to it, is that we want to make sure that when we update each of the components of the model, we don't forget about the knowledge that is stored in their parameters. And So what is the idea there is that what we do is we know that there are like multiple MMP blocks, each of them are updated with different frequencies. And then again, the fast and slow here is just a relative term. It doesn't mean that you know the slowest or the fastest. So we have a slow and fast way. It can be any part of the neural network and we want to transfer the knowledge from one to the other. So in order to do that, what we do is is a distillation process that I mentioned, but the distillation process is based on the on policy distillation. So the model itself generates some data. And so one interpretation of of this process is that we distill the knowledge of one small model to a larger model. When we want to go from one step to the next one, we activate new parameters in the next level. So it can help the model to release some of its capacity and be ready to accept new knowledge. So somehow it is somehow it describes a very natural way of learning in humans as well. And for example, it's, it's very common that when we learn something, when a new, when, when we learn about the new concept or something like that, we don't have a full understanding of all asects of that. But the point is when time passes and you know, we, we let the, we let our brain to better understand that concepts overtime. And also like, you know, we, we study other things and better understand the entire process. Then at some point we can see that we have a very clear picture of what's going on in that concept. We can completely understand it. And so generally that's the best thing. erstand the entire process. Then at some point we can see that we have a very clear picture of what's going on in that concept. We can completely understand it. And so generally that's the best thing. That's a very good way of learning. So here is, is exactly the same, very similar process, very, very similar process. So also like discuss that from this perspective, I think it might be better and more simple way to understand why we need like have multiple levels and distill the knowledge from each level to adopt another one. So when we want to understand a specific concept, we have different levels of knowledge abstraction for ourselves. The first level which is the most simplest 1 is to just memorize things. So let's say that we want to learn one specific mathematical rule or for example, a specific concept in physics or any science. How how we can learn that? We start with some example of that specific concept and then we start memorizing those concepts. For example, if if it's a mathematical rule saying that you know, just any mathematical rule that we can have, we start with some specific examples and just memorize them. And then at some point, we just neuralize our understanding of all those examples, remove all those examples in our brain, and replace all of those memories with just one single memory that can describe everything that we have learned so far from that concept. And then when time passes, we have more information, we read more about that concept and so on and so forth. Again, we revisit our understanding of the concept and then replace our previous understanding with this nuke understanding, which is more general and it can explain a more phenomenal or more terms in that specific concept. So that's generally the way we understand things and there are different levels of abstraction in our understanding. So now when we have different MLP blocks or or generally, let's just go to any like arbitrary architecture, There's the don't, it doesn't have to be just hope. It can be any architecture. But the, the main thing is that each of them are, each of the blocks are updated with different frequency. In that case, the fast updating block is very similar to the memorization process because we memorize a lot of things. We don't need to understand that. There's no like pure understanding of that concept. It's just memorization. And we can also like forget very fast something that we have memorized. So the first level is that. So the first block is responsible for that part. But if we want to better understand that concept, we need to do some memory consolidation. So what is happening in our design is that we transfer the knowledge from fast updating module to the other one. But if we just simply pass the knowledge from fast updating block to the slow updating block, then nothing has changed. We just like transfer the knowledge without doing anything. le to the other one. But if we just simply pass the knowledge from fast updating block to the slow updating block, then nothing has changed. We just like transfer the knowledge without doing anything. But instead of just simple transfer, we replace that with the solution process. Why this solution here is important? Because the previous block or generally the fast updating block has compressed the concept and somehow understand it or memorize it in any way. Just it's just a compression process. When we do a solution then there is another levels of compression that force the model to, you know you don't have that all those parameters anymore. You have now have less number of parameters to store that specific knowledge and in order to do that you need to come up with something that is more general and can can understand underlying patterns in the data in a better way. So you can store everything in just a smaller number of parameters. So in that case, the model would come up with better levels of knowledge abstraction because we have forced it to do it. And then again, we just repeat this process so on and so forth. That's generally the main idea of memory consolidation. And like every time that that this sleep process happens, we consolidate the knowledge from one level to the other one and so on and so forth. And so that's, that's a very high level idea of what's happening in the memory consolidation. Another part is about dreaming. So why we need to have this dreaming process? The main thing is I I think there are like 2 important points when everyone wants to implement this streaming process. The first part is we need to have a self improvement process. So we have learned something so far. Actually the the memory consolation part can also be seen as a form of self improvements. But you know, if we haven't won a specific task at hand, if we want to like a specifically optimize the model for one task, then this is the place that we can do it. We can self modify the model and like fine tune it, or generally like use RL to update the models and self modify it so it can be more powerful in one specific task and so on so forth. That's just one advantage of trimming. Another advantage of trimming is that in the dreaming process we need to understand the connection of concepts that seems to be irrelevant, but they are actually relevant. That's also what is happening in the dreaming process of human We can see that we can. We can see very weird dreams because the brain is trying to understand the connection of of very irrelevant concepts and see whether there is an underlying pattern in that O. Here in the dreaming process we need to also have that as well and understand different aspects of how we need to combine different knowledge stored in different components of the model. So that's another goal of the dreaming. And you know, we can just combine these two into the sleep process and the model after one step of sleep, the model has consolidated its own memory. And also on the other hand, there's a self improving process.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence