Evidence receipt / evaluation
Published · transcript-backedAli Behrouz: evaluation
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability.”
Source trail
Everything needed to verify it.
- Speaker
- Ali Behrouz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, I think the key phrase maybe for people to latch onto there, starting with myself is modifying its own update rule. And this is definitely a theme in general, like with the Mamba architecture as well. You know, the authors of the Mamba architecture had done a bunch of previous state space model work and the big unlock at the with Mamba specifically was that the way in which the state is going to be updated at each time step now became a function of inputs. And the so that increased expressivity potential for expressivity of making the the actual update of the state itself dependent on the input that it's receiving at that time unlocked a lot better performance. And there's something very similar going on here it seems like where you're saying we want to make the final value output, not some. We don't want to, we don't want to calculate the value vector too early. Basically we want that to come a little bit later and be more input dependent, more history dependent than it has has traditionally been in the softmax attention. Is that a good intuition or is there is there something still missing from that intuition I. I, I think it's, it's, it's a perfect intuition. Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability. Generally the model itself is also very adaptive to the context. And so from every token that comes, the model is trying to learn something. And so the way that it generates the value from the value term for the associative memory is exactly the same as the way it updates its memory. So it's it's a very adoptive process to also generate the value. So can you take us through kind of a single time step and maybe we can do this with the attention hope and then the Titans hope and just like highlight the little difference there, but it's also zoom out to the big picture. So we've got now kind of a new fundamental block, right? If we do the attention hope thing, first, we've got an attention mechanism and then we've got multiple MLP's arranged in sequence from fastest update to slowest update frequency. And then that block gets stacked into layers, correct? I'm a little bit confused to be honest. I'm the sort of there's no training test distinction because there is still some training process where you're like just taking a bunch of data and like running it through the thing, right? So from the sort of researchers perspective, maybe from the models perspective, there's not so much of A distinction, but from the researcher perspective, you are still sitting there like running a process that takes a bunch of data and like has the model learn from that, which is kind of, you know, a bulk process that's not like a user is engaging every there's nothing like outside of that process happening at that time, right? So what, how is it different? You know, when I, when I do this for a transformer, I can like, I do have this some of these really nice parallelization benefits. I am interested to come back to understand like to what degree the sort of current hardware paradigm plays nicely with some of the stuff you've got here and to what degree the, you know, the fundamental recurrence may present challenges, but bracket that for a second. Today I can like run a bunch of tokens through the thing in parallel. We can accumulate, you know, we can do this for batches. We can accumulate all these gradients and then we sort of apply the gradients and then we have, you know, the the next time stamp and we kind of keep doing that. And I have a pretty good intuition for like how information flows. I kind of, you know, can I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and, you know, gradually updating all the weights. How does the procedure with the new architecture vary? What are the core things that are different from the the paradigm that we're more used to, I think.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.