Evidence receipt / belief
Published · transcript-backedAli Behrouz: belief
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“Do you want to use existing pre trained model or do you want to like start from scratch and train your own designer architecture? But I think potentially both of them are relatively similar.”
Source trail
Everything needed to verify it.
- Speaker
- Ali Behrouz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Did I have it right that there are like this? This core block of either the traditional attention or the self modifying Titan module plus the MLPS that then becomes the block that gets stacked into layers? Is that right? Yes, that's that's a design choice. We can have like different design choices. For example when we. So generally the initial and and the main design of Hope is in the case that we have. For example for the Hope attention, we have attention and then multiple MLP blocks. Each of them are updated with different frequency. But for some of the task we needed to use pre trained model and for example if if we want to focus on for example Llama, then Llama is not designed with Hope architecture. So what we can do, the thing that we have done is that instead of like going in this formulation that I mentioned, for example, attention then multiple MLP blocks, what we have done is that we say this is attention and MLP block, then attention and another MLP block with different frequency and then attention and another different, sorry, another MLP block with different frequency and so on so forth. So it's somehow design choice you need to see like which one do you prefer? Do you want to use existing pre trained model or do you want to like start from scratch and train your own designer architecture? But I think potentially both of them are relatively similar. They don't fundamental, they make changes, yeah. Interesting. It's some of the, I mean, great reminder of the old Ilia maxim that the models just want to learn. There's a lot of, it's always striking to me how many of these choices end up kind of, yeah, I could kind of go one way or the other. I mean, that was true in the Titans case where you had like 3 different ways of working the memory module into the larger architecture. Definitely been true with Mamba in many ways where you know, you could have multiple states, you can have them in sequence, you can have them in parallel, You can, you know, once you have one of these block concepts that seems to work well, you can kind of Lego piece it in a lot of different ways. And yes, there will probably be some performance differences between different ways to arrange the blocks. But more often than not, it seems like if you're really talking about a a serious conceptual advance, you find that like that's less important. The exact wiring diagram is less important and more important is the core piece that you're adding to the the set of Lego pieces sort of speak that you can use. And so, yeah, that's a good reminder of that. How do you think about the relationship between the different MLP's in terms of size, in terms of update frequency, maybe in terms of like learning rate? I, I because it feels like there might be some feels like learning rate might be kind of important here where there's like a potentially an equivalence. If I update 1 LLP every token and then I update another one with a larger batch size, I can, I feel like I can make those like much more similar or quite a bit different depending on like what learning rate I apply. But yeah, so size frequency like intangible terms. Are the small, are the ones that are updated less frequently also smaller and do they have like different learning rates than the other ones? Like take us through that kind of how do you think about the the relationships between the MLPS of different?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.