High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“The core idea is I see it in the nested learning paradigm. And what I, what I think is like potentially for simple person like myself, like most exciting about it is for quite some time now, right, we have achieved the greater and greater expressivity of models by stacking more and more layers and just making them bigger.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…erlying patterns between the received data. So that's a pretty high level of of inspiration from the brain. And so, yeah. And in short, I, I, I think that we don't want to replicate what human can do. And you don't want to have human intelligence. But on the other hand, we want to have a new form of intelligence that is really, you know, it's really proper and it's in there, you know, it's, it's designed in a good way that understands human needs. And so it can help people to, you know, do a lot of things that they might face some challenges without illness. Yeah, certainly, they're already superhuman in some ways. And so the opportunity for them to balance out our weaknesses is incredible. 1 phrase that you said there that I wanted to kind of latch onto is multiple ways to define intelligence. And maybe I'll just give you kind of my high level pitch for like what nested learning is. I think in the paper, there is a big emphasis, you guys place a big emphasis on kind of showing certain equivalences where you're like the way that we're doing things today is sort of a special case of a more general framework that you're developing. The core idea is I see it in the nested learning paradigm. And what I, what I think is like potentially for simple person like myself, like most exciting about it is for quite some time now, right, we have achieved the greater and greater expressivity of models by stacking more and more layers and just making them bigger. And that has worked like remarkably well. We've been able to push that that paradigm incredibly far. It's like kind of crazy though, right, That we just have this like one layer stacked over and over and over again in, you know, you know, whatever 80 or 120 layers, zebra, however many layers. And that's kind of it like it, that feels like the more mature solution should be somehow more like elaborate than that, right? And so, but that's been the way that we've achieved this expressivity, or sometimes the term computational depth is thrown around and what necessary learning is doing is bringing a different way to the table to achieve higher levels of exclusivity or or higher levels of computational depth. And that is by stacking not layers, but levels. And what differentiates a level from a layer is that a layer is like the same thing. Or if they can, they can alternate. Obviously we have these sort of, you know, interleaved architectures too, but these are things that are sort of in sequence as information passes from one layer to the next. But kind of a forward pass is like pass through all the layers 1 by 1 and get to the end. And that's kind of the thing what the levels paradigm brings to it that's different is that different levels can have different update frequencies. And with that, you now have the possibility for some parts of the overall system to be much more durable and some parts to be updating like much more radically in something much closer to real time. And that obviously feels like just much more aligned to like what we are, right? We're not like one static thing that processes information in a fully dependent way each time we have, you know, we're, we're very much our state in any given time is very much contingent on what we just experienced, but only to a degree, right? Like where, you know, my mood or what it, what is currently on my mind is a reflection of what happened earlier today. But my like big picture views about the world, you know, they didn't change from this morning until now. y mood or what it, what is currently on my mind is a reflection of what happened earlier today. But my like big picture views about the world, you know, they didn't change from this morning until now. So there's clearly some sort of hierarchy of different kinds of beliefs, different kinds of representations, different kinds of circuits that we have, which are updated in some cases very quickly and other cases very slowly. And obviously they like are integrated together and work together. And we just haven't seen that in machine learning, except maybe in a few, you know, very far-flung kind of experimental cases. And now you're like really starting to show that with this nested learning paradigm, not only can you make it work, but as we get into with results, like you can make it work in a way that is competitive with Transformers and even seems to have some of these new or sometimes called them micro skill advantages where you can see, you know, with these certain early diagnostics that like, oh, this can do something that's like qualitatively different than what a transformer can do, even as it's also like outperforms it a bit in terms of like general perplexity type scoring. So how would you react to that kind of general, you know, high level summary? And then I also really would be interested to get your take on what is this concept of computational depth or expressivity. I'm tempted in some ways to make an analogy to just like the G factor. You know, people talk about the obviously the G in AGI, it is the generality. There's also G in the context of like human IQ, which is like the G factor of the sort of intangible something that's like how you know, how capable are you across like a very wide range of things. Again, that's like just getting it in generality. Maybe in machine learning it's as simple as being like G is sort of loss or there's maybe some fundamental equivalence there, but maybe not, I don't really know. So I'm very interested in how you think about what that clearly we're getting at something that we've seen like huge progress, but what is that something is another thing I really would love to get your intuition on.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence