Evidence receipt / observation
Published · transcript-backedAli Behrouz: observation
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“The you know, the model can you know the first MLP block can forget something. But the point is if that specific data sample or that specific skill that is forgotten is important, then this can come back to the other MLP blocks which has not been updated so far and they have still they have the knowledge about that specific skill or data sample.”
Source trail
Everything needed to verify it.
- Speaker
- Ali Behrouz
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Can you just describe in more specific detail, like what are the relative sizes of the levels? What are the structures of the levels? What are the context windows or lengths of the different levels? What are the frequencies? Just like map the thing out for us in kind of very black and white terms. Yeah. So and we start from the Transformer structure. In Transformer we have attention block and then MLK block. So what is happening there is that in the pre training we have you know different contexts that attention is attent to. The attention side is trying to like combine all the tokens and each token attends to all the tokens before that in the context and so on. So first and then there is an MLP block and that MLP block is responsible for long term memory. Now when the model is pre trained then the MLP block is fixed, it's not changing anymore. So it has all the information compressed during the pre training and then we have attention. So in at inference time, attention is responsible for the context that is it is getting and MLC block is responsible for the long term memory and generally the very general knowledge of the world. Now let's just simply extend this idea. A simple extension of this idea is that we can simply keep attention and then instead of just one MLC block, we have multiple MFV blocks. Each of them are updated with different frequency. So now what is going on there? Is that why it's helpful is that when you have attention, you have a fast adoption to the context. Attention is very powerful. It's like a perfect memory cache everything. And so it's it's great. On the other hand, you might want to have multiple levels of memory and that's the part we define continuum memory system. So instead of 1 block of MLP, we have multiple blocks of MLP. And now you have your first MLP block, it is updated very fast. So what would happen in that case? In that case, the first updating process of this MLP block can cause catastrophic forgetting, because this MLP block can simply forget the information that it it gets, for example a couple of thousands tokens ago, for example. But the point is since the other MLP blocks have not updated so far, the knowledge that is forgotten by the you know first MLP block is still in their parameters. So when we perform back propagation through all these layers, then the knowledge can come back. So it provides and it helps us to have a loop process in time. The you know, the model can you know the first MLP block can forget something. But the point is if that specific data sample or that specific skill that is forgotten is important, then this can come back to the other MLP blocks which has not been updated so far and they have still they have the knowledge about that specific skill or data sample. So that's a very simple way of extending whole transformer block, transformer block. And so we call this variant hope attention. You know, it's, it's a combination of attention plus multiple MLP blocks. And that's what we called hope attention. Now we have another variance, which is the actual hope architecture. What we are saying is that attention, as I mentioned, is as a perfect memory. It's, you know, it can cache everything, it can scale and it's fast and so on so forth. It's, it's great. pe architecture. What we are saying is that attention, as I mentioned, is as a perfect memory. It's, you know, it can cache everything, it can scale and it's fast and so on so forth. It's, it's great. But the point is, still the update of attention has infinite frequency. What does it mean? It means that Attention doesn't know anything about the temporal dependency of all the tokens and it needs something like positional encoding. Or even with the help of positional encoding, Attention is not a great model for the tasks that are sequential task that requires sequential reasoning or something like that. What was our idea? Our idea was to replace attention with another associative memory that's map keys to values. So that's that's what attention does. And now we want to replace it with another module that tries to map keys to values. And so one potential architecture here is title. So we can simply just replace Titan and have Titan plus continue memory system. And so that's that's a simple idea. But the point is, you know, in initial sections of the paper we discussed that if you have a simple linear process of updates, which is what is happening inside each chunk of Titan update, this process can somehow be. Be richer than the case, then we have self referential process. So what is self referential process? Gradient descent or generally back propagation is a form of self referential process. So what is the idea in the self referential process? The idea there is that we want to learn how to learn and how to learn and how to learn and how. There are a lot of levels of how to learn how to learn. And so computationally it's it's infeasible to implement all those levels of how to learn how to learn and so on so forth. So there is one idea, but by Schmidhaber ET al. And so basically they have this idea of self referential model and how, for example, one of the ways that we can make a model self refresher when we have a key value memory is the case that the model generates its own value. So what is happening there? Let's say that in our memory we want to memorize something and you know there is one specific event happening and you want to memorize it. Let's say you want to memorize one specific word. And now in our brain we have associative memory and we are trying to, you know, map this specific word to another concept that we already knew so we can memorize it. And you definitely have seen some cases that, for example, you say I generate this specific word or I, I want to like map this specific word to another word that I already knew. So I can remember that we, we generate the value that we want to map our keys to it. So we can, you know, memorize the key as well. So for the set referential process, it's a very general concept, but in this specific design choice that we have in the paper, the value of the associative memory is generated by its own parameters. So the model itself generate its own value and then try to map keys to values. So potentially this process is fully sequential. You cannot parallelize it in a simple format. And so it has a full understanding of the causality in our data.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.