Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Atom standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again, the argument is that it more than pays back for itself in terms of faster convergence or just better learning.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Math data is also like. It's very similar to the recall intensive tasks. But the point here is that you know there are different setups for it. For example, in one of them is the noisy in context recall. We want to perform recall and in context recall task. But the point is we have some noise in the tokens and now when we have that noise in the tokens, somehow the power of Transformers that I explained in the previous setup which is which was pure in context learning. Now is its weakness somehow, because it can get simply confused about which token is noise, which token is not and so first. So potentially this task becomes a little bit, for example, harder for Transformer compared to a model like whole. But again, that that is also very that also depends on the memory management of the RNN. So again, for the RNN, if it doesn't, if it doesn't have a very good memory management system or generally updates mechanism, then potentially it can simply get confused by the noise as well and face some issues. But again, if the memory management is strong, then it's, it's much simpler to filter all those noise tokens in the task. And so yeah, I, I think that's for example, one thing, another task that is also like, I think it's interesting here is about compression. So the, the compression task here, somehow you know, the, the name explain the task itself, but we want to compress the tokens and predict one single token that is the compressed version of a set of tokens. And so then, you know, we see that and then we want to like reconstruct the original sequence from that. And so potentially it's a simpler task for models like RNN because they already knew how to compress the data properly. But on the other hand, Transformer has has a harder time to perform this task. So generally, like as I mentioned, all of these tasks I don't want to like go into all of the are of the research. But generally all these tasks are somehow modified version of recall intensive tasks or in context recall or something like that. But the point is, there are other aspects. For example, selective coping is another one. There are other aspects to the model that those aspects are so are very important and we should also like see how the model performs in those aspects and not just, you know, over freeze our evaluation on one specific metric. Cool. I think that's probably enough on the really low level stuff. And I think this illusion of architecture title of the paper starts to click for me. On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Atom standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again, the argument is that it more than pays back for itself in terms of faster convergence or just better learning. Is there anything you want to add on the the M3 optimizer as you call it? First, also like could I have one point that generally for optimizers it's a little bit like hard to say that for example, this specific optimizer is more powerful than the other one. It really depends on the problem setup or generally even the problem. So for example, we might here we are like evaluating the optimizer on, for example, vision task. But on the other hand, if you train a language models, you might see that the trend is completely different or something like that. So generally like the design of optimizer and and saying that like which one is better than the other one really depends on the task. Or if you take share problems, it's up and all these things. And somehow that's also one of the main points that you wanted to deliver in this learning. Because what we are saying is that the entire architecture with its optimization process are just one interconnected system of nested optimization problem. And this is interconnected. Why it's interconnected? Because the gradients of the optimization side is generated by the architecture. If you have a simple architecture, then the gradients are very simple. If you have a complicated architecture, the patterns in the gradients can be very complicated. And then when you have momentum, term momentum is a form of is a form of associative memory that is trying to compress gradients. So for example, if if your gradients are very complicated, you need more powerful memory management system for your momentum. Or if if it's very like simple, if it's very simple architecture, then even a simple gradient descent without any momentum might work very well. So in general, one of the arguments that we have in the paper is that we should see everything as an interconnected system and tries to design something that's all together results in a good model architecture or generally like a machine learning model in a very general term. So that's that's one argument. Another thing is that we wanted to deliver this message that architecture site is very, very, very similar or somehow exactly the same as optimization side. All of them are just some learning room and there are some learning process that is happening. And the only difference between the architecture side and the optimization side is just the context. The context of the optimization algorithm is gradients. Actually, the context is the set of gradients that we have. And the context of the architecture side is a set of tokens that we have. So generally they are very similar. So in the paper we had this continue memory system, we extend the MFP block saying that you can have multiple levels of frequency for the MFP block. And you know that's a very general term. And in the entire paper we are arguing that architectures are the same as optimizers and so on and so forth. So why not applying that technique and somehow borrow that technique from architecture side and apply it to the optimization side.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.