Evidence receipt / observation
Published · transcript-backedAli Behrouz: observation
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“We want to perform recall and in context recall task. But the point is we have some noise in the tokens and now when we have that noise in the tokens, somehow the power of Transformers that I explained in the previous setup which is which was pure in context learning.”
Source trail
Everything needed to verify it.
- Speaker
- Ali Behrouz
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…And what does the mad data set get at? How should we understand like because that's, that's where to just replay back to what you just said. The on these like needle in a haystack. These like very difficult recall tasks from earlier in context, the transformer remains the best. The recurrent models, which only have some latent representation and and don't have the ability to look back at the original raw text don't perform as well. But with each generation of improvement, and here you've got several, the hope architecture does the best of of the recurrent ones that don't have the the full explicit context in working memory at runtime. And then so that gap is closing, flipping over to the mad data set. Here the hope architecture is performing better than everything including the transformer. What like micro skills, is that testing which? Which should we take away from that result? The. Math data is also like. It's very similar to the recall intensive tasks. But the point here is that you know there are different setups for it. For example, in one of them is the noisy in context recall. We want to perform recall and in context recall task. But the point is we have some noise in the tokens and now when we have that noise in the tokens, somehow the power of Transformers that I explained in the previous setup which is which was pure in context learning. Now is its weakness somehow, because it can get simply confused about which token is noise, which token is not and so first. So potentially this task becomes a little bit, for example, harder for Transformer compared to a model like whole. But again, that that is also very that also depends on the memory management of the RNN. So again, for the RNN, if it doesn't, if it doesn't have a very good memory management system or generally updates mechanism, then potentially it can simply get confused by the noise as well and face some issues. But again, if the memory management is strong, then it's, it's much simpler to filter all those noise tokens in the task. And so yeah, I, I think that's for example, one thing, another task that is also like, I think it's interesting here is about compression. So the, the compression task here, somehow you know, the, the name explain the task itself, but we want to compress the tokens and predict one single token that is the compressed version of a set of tokens. And so then, you know, we see that and then we want to like reconstruct the original sequence from that. And so potentially it's a simpler task for models like RNN because they already knew how to compress the data properly. But on the other hand, Transformer has has a harder time to perform this task. So generally, like as I mentioned, all of these tasks I don't want to like go into all of the are of the research. But generally all these tasks are somehow modified version of recall intensive tasks or in context recall or something like that. But the point is, there are other aspects. For example, selective coping is another one. There are other aspects to the model that those aspects are so are very important and we should also like see how the model performs in those aspects and not just, you know, over freeze our evaluation on one specific metric. Cool. I think that's probably enough on the really low level stuff. And I think this illusion of architecture title of the paper starts to click for me. On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Atom standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again, the argument is that it more than pays back for itself in terms of faster convergence or just better learning. Is there anything you want to add on the the M3 optimizer as you call it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.