High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Ali Behrouz: evaluation

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“Now we just need to understand how what we have been already doing is, is a form of in context learning. And so that was a part we started to like showing that for example, back propagation is a form of in context learning is a form of associative memory.”

— Ali Behrouz

Source trail

Everything needed to verify it.

Speaker
Ali Behrouz
Attribution
Verified speaker
Claim type
evaluation
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…You know. Combine or mix the data and something like that. In that case, I can see that the quality can potentially goes up, the quality of next token prediction for this specific design choice can goes up. And why is that? Because it can be interpreted as as as a form of internal thinking process. So now for each specific parameter that I have in my model, it is performing favor all computation. So it's not just one parameter 1 of computation, it's one parameter couple of steps of computation. So that's one advantageous. Another one is about memory perspective and adoption of these models. So when we have a model that is like adopting to the context very fast, then potentially that model can, you know, learn in context. So that's that's somehow one of the main messages that we try to deliver in the nest of learning, which was everything that we know of somehow is a form of in context learning. So generally, like, I think it's a great thing in human language that we create new words. But on the other hand, if we just create a lot of words for the same concept, it can just make us confused or or it can help, it can be misleading somehow. So I think we should like create new words to differentiate different concepts. But if we have one specific concept, we need to like stick to one specific word that we have for that concept to avoid misleading process or anything like that. So from that perspective, we realize that we can say everything is just a form of in context learning. So we already know that's what is in context learning. Now we just need to understand how what we have been already doing is, is a form of in context learning. And so that was a part we started to like showing that for example, back propagation is a form of in context learning is a form of associative memory. And when it's a form of associative memory, then we can say that the general pre training phase of the model is a form of in context learning. Or when we go to the for example, context of attention or RNA, so on and so forth. Again, it's a form of in context learning. So when we perform gradients which we can define any RNN based on gradient descent or or other form of optimization process, when we can do that, it means that we are doing some learning on the context that is happening right now. So I think these two are the main things that nested learning is trying to address. 1 is about generally more computation per per neuro and another one is about adoptable tea and continue all their visa. Can you just describe in more specific detail, like what are the relative sizes of the levels? What are the structures of the levels? What are the context windows or lengths of the different levels? What are the frequencies? Just like map the thing out for us in kind of very black and white terms.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence