High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

The Cognitive Revolution / episode intelligence

Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

3 Jun 2026 22 published claims 2 attributable people

Speakers in the public record

Claim mix

belief 9evaluation 4observation 2uncertainty 2prediction 2recommendation 2preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

22 published records

01 / belief

I kind of, you know, can I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and, you know, gradually updating all the weights.

“I kind of, you know, can I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and, you know, gradually updating all the weights.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

02 / belief

Like different aspects to answer this question from the technical point of view, I I think generally the entire literature, most part of the literature in the, you know, past 40 years is built on a paradigm that says we have a pre training or generally training phase and we have a test phase.

“Like different aspects to answer this question from the technical point of view, I I think generally the entire literature, most part of the literature in the, you know, past 40 years is built on a paradigm that says we have a pre training or generally training phase and we have a test phase.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

03 / belief

The core idea is I see it in the nested learning paradigm. And what I, what I think is like potentially for simple person like myself, like most exciting about it is for quite some time now, right, we have achieved the greater and greater expressivity of models by stacking more and more layers and just making them bigger.

“The core idea is I see it in the nested learning paradigm. And what I, what I think is like potentially for simple person like myself, like most exciting about it is for quite some time now, right, we have achieved the greater and greater expressivity of models by stacking more and more layers and just making them bigger.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

04 / belief

On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Atom standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again, the argument is that it more than pays back for itself in terms of faster convergence or just better learning.

“On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Atom standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again, the argument is that it more than pays back for itself in terms of faster convergence or just better learning.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

05 / belief

I mean, I was, I always say the last and least valuable co-author of the emergent misalignment paper that came out about a year ago and there's been a lot of variations on that since. But the kind of big take away the big theme, you know that I think we should all we would all do well to remember is changes to, I don't want to say 1 area, but sort of changes made to a neural network with one particular purpose or one particular data set can have like very strange and surprising knock on effects in behaviors that at first glance would seem like very far afield.

“I mean, I was, I always say the last and least valuable co-author of the emergent misalignment paper that came out about a year ago and there's been a lot of variations on that since. But the kind of big take away the big theme, you know that I think we should all we would all do well to remember is changes to, I don't want to say 1 area, but sort of changes made to a neural network with one particular purpose or one particular data set can have like very strange and surprising knock on effects in behaviors that at first glance would seem like very far afield.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

06 / observation

We want to perform recall and in context recall task. But the point is we have some noise in the tokens and now when we have that noise in the tokens, somehow the power of Transformers that I explained in the previous setup which is which was pure in context learning.

“We want to perform recall and in context recall task. But the point is we have some noise in the tokens and now when we have that noise in the tokens, somehow the power of Transformers that I explained in the previous setup which is which was pure in context learning.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

07 / belief

I mean, maybe it's been somewhat fruitful, but I think also people would get very confused and they try to interpret dreams or understand, you know, what's going on there.

“I mean, maybe it's been somewhat fruitful, but I think also people would get very confused and they try to interpret dreams or understand, you know, what's going on there.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

08 / uncertainty

I might be wrong, but I think the least level of criteria that that we could say something is conscious is that is that that model or that that being be active, it has a form of active processing the information.

“I might be wrong, but I think the least level of criteria that that we could say something is conscious is that is that that model or that that being be active, it has a form of active processing the information.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

09 / belief

I think I want to spend one more beat on what you mean when you say generating its own value, because I'm kind of like, OK, I, I know the transformer architecture pretty well.

“I think I want to spend one more beat on what you mean when you say generating its own value, because I'm kind of like, OK, I, I know the transformer architecture pretty well.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

10 / belief

Do you want to use existing pre trained model or do you want to like start from scratch and train your own designer architecture? But I think potentially both of them are relatively similar.

“Do you want to use existing pre trained model or do you want to like start from scratch and train your own designer architecture? But I think potentially both of them are relatively similar.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

11 / belief

You know, generally also there are some challenges in the infrastructure of of the model for more board modelling. And so all these things together, I think there are more important, you know, tasks for, for example, word modeling rather than start starting to work on these specific design shoes, but differently at some point when we could solve all those challenges, then then we can come back and use all these techniques for free, further improving.

“You know, generally also there are some challenges in the infrastructure of of the model for more board modelling. And so all these things together, I think there are more important, you know, tasks for, for example, word modeling rather than start starting to work on these specific design shoes, but differently at some point when we could solve all those challenges, then then we can come back and use all these techniques for free, further improving.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

12 / observation

The you know, the model can you know the first MLP block can forget something. But the point is if that specific data sample or that specific skill that is forgotten is important, then this can come back to the other MLP blocks which has not been updated so far and they have still they have the knowledge about that specific skill or data sample.

“The you know, the model can you know the first MLP block can forget something. But the point is if that specific data sample or that specific skill that is forgotten is important, then this can come back to the other MLP blocks which has not been updated so far and they have still they have the knowledge about that specific skill or data sample.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

13 / preference

When we are talking about in context recall tasks or or generally like recall intensive tasks, in my opinion all those tasks are designed for Transformers.

“When we are talking about in context recall tasks or or generally like recall intensive tasks, in my opinion all those tasks are designed for Transformers.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

14 / evaluation

Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability.

“Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

15 / uncertainty

Because I might think, jeez, you know, all the stuff I've done with this model in this One Direction, like probably isn't going to help me over here.

“Because I might think, jeez, you know, all the stuff I've done with this model in this One Direction, like probably isn't going to help me over here.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

17 / prediction

I I think even from the research point of view, it might be better to say that we have evaluation time and not evaluation time because it seems we are always like update when when the model is always up, gets updated over time.

“I I think even from the research point of view, it might be better to say that we have evaluation time and not evaluation time because it seems we are always like update when when the model is always up, gets updated over time.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

18 / recommendation

Avoid such cases because when we are in this context, let's say that, for example, let's say that I, I don't know anything about one specific task.

“Avoid such cases because when we are in this context, let's say that, for example, let's say that I, I don't know anything about one specific task.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

19 / evaluation

For example you can or when you are doing deep learning any form of deep learning and you are saying that I'm using this attention here you are actually using nested learning.

“For example you can or when you are doing deep learning any form of deep learning and you are saying that I'm using this attention here you are actually using nested learning.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

20 / recommendation

Even you say, for example, write this e-mail for me, they can simply write that for you. And so I think that's a very promising direction in general, because we don't want to replicate human.

“Even you say, for example, write this e-mail for me, they can simply write that for you. And so I think that's a very promising direction in general, because we don't want to replicate human.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

21 / evaluation

For example, if if it's a mathematical rule saying that you know, just any mathematical rule that we can have, we start with some specific examples and just memorize them. And then at some point, we just neuralize our understanding of all those examples, remove all those examples in our brain, and replace all of those memories with just one single memory that can describe everything that we have learned so far from that concept.

“For example, if if it's a mathematical rule saying that you know, just any mathematical rule that we can have, we start with some specific examples and just memorize them. And then at some point, we just neuralize our understanding of all those examples, remove all those examples in our brain, and replace all of those memories with just one single memory that can describe everything that we have learned so far from that concept.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution

22 / evaluation

Now we just need to understand how what we have been already doing is, is a form of in context learning. And so that was a part we started to like showing that for example, back propagation is a form of in context learning is a form of associative memory.

“Now we just need to understand how what we have been already doing is, is a form of in context learning. And so that was a part we started to like showing that for example, back propagation is a form of in context learning is a form of associative memory.”
Speaker
Ali Behrouz
Publisher
The Cognitive Revolution
Search evidence