High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“I kind of, you know, can I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and, you know, gradually updating all the weights.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…I, I think it's, it's, it's a perfect intuition. Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability. Generally the model itself is also very adaptive to the context. And so from every token that comes, the model is trying to learn something. And so the way that it generates the value from the value term for the associative memory is exactly the same as the way it updates its memory. So it's it's a very adoptive process to also generate the value. So can you take us through kind of a single time step and maybe we can do this with the attention hope and then the Titans hope and just like highlight the little difference there, but it's also zoom out to the big picture. So we've got now kind of a new fundamental block, right? If we do the attention hope thing, first, we've got an attention mechanism and then we've got multiple MLP's arranged in sequence from fastest update to slowest update frequency. And then that block gets stacked into layers, correct? I'm a little bit confused to be honest. I'm the sort of there's no training test distinction because there is still some training process where you're like just taking a bunch of data and like running it through the thing, right? So from the sort of researchers perspective, maybe from the models perspective, there's not so much of A distinction, but from the researcher perspective, you are still sitting there like running a process that takes a bunch of data and like has the model learn from that, which is kind of, you know, a bulk process that's not like a user is engaging every there's nothing like outside of that process happening at that time, right? So what, how is it different? You know, when I, when I do this for a transformer, I can like, I do have this some of these really nice parallelization benefits. I am interested to come back to understand like to what degree the sort of current hardware paradigm plays nicely with some of the stuff you've got here and to what degree the, you know, the fundamental recurrence may present challenges, but bracket that for a second. Today I can like run a bunch of tokens through the thing in parallel. We can accumulate, you know, we can do this for batches. We can accumulate all these gradients and then we sort of apply the gradients and then we have, you know, the the next time stamp and we kind of keep doing that. And I have a pretty good intuition for like how information flows. I kind of, you know, can I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and, you know, gradually updating all the weights. How does the procedure with the new architecture vary? What are the core things that are different from the the paradigm that we're more used to, I think. Generally like the main difference comes from the update side that I mentioned. So for the Copa attention, I think it's very similar to the current paradigm and it's the even the architecture is very similar to Transformers. It's actual transformer architecture. We just replace the MLP block with multiple MLP blocks. And so I think when we want to do inference, then the main difference comes from the fact that for each of the MLP blocks, we need to track where we are. Is it the time that we want to update the MLP block or it's it's a still test time get updated. If it's the later case, then we use the last updated state of the that specific MLP for for doing the inference. If it has not, I mean, if it's the time to get updated, then we first update it through all the back propagation stuff and all of the tokens that we have seen so far in the in the current chunk. And then when, when the weight is updated, then we perform the, you know, inference. From the research point of view, it might be a little bit hard to remove this part. I I think even from the research point of view, it might be better to say that we have evaluation time and not evaluation time because it seems we are always like update when when the model is always up, gets updated over time. There is no training time and test time. But the point is it seems that for, you know, one specific period of time, we don't do some evaluation and we wait for some time. And after that we start evaluating the model on the different downstream tasks that we have. Generally, like for anything else, I think from the model perspective, it doesn't know whether it is in the test time or train time because everything is the same and it's a very uniform process. But from our side, definitely it is important whether we want to evaluate the model and you know, measured its accuracy for a specific task and so on. So first or not? And so, yeah, coming back to the hope architecture, that's generally for the for the hope transformer or Hope attention model that I mentioned. But when we go to the the actual hope architecture, then again, everything is the same, Everything is very similar. The only difference is that the attention is replaced by self modifying Titan and for self modifying Titan, again, everything is very similar to Titan. So the are, all of the are are just inside the model design. And from the higher level perspective the inference is very similar that you know the context or document goes to the self modifying Titan and then for each token we have one output and then it goes to the MLP blocks, we have multiple MLP blocks and so on and so forth. So everything is is very similar to the current paradigm.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence