Evidence receipt / uncertainty
Published · transcript-backedNathan Labenz: uncertainty
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“Because I might think, jeez, you know, all the stuff I've done with this model in this One Direction, like probably isn't going to help me over here.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Any evaluation that we have used for like Hope architecture in Nested learning paper potentially can be done here as well. So then the goal is exactly the same thing. At the end of the day, the model needs to continually learn new knowledge and new, you know, learn about new tasks, learn about new skills and so on and so forth. So in some sense the goal is very similar, but I think the set up of the problem is the part that is different from this paper and this learning. In asset learning, we are talking about the active phase of the model, but here we are talking about the sleep time of the model. So that's generally the main difference. But all of the evaluations can be done and we can we can see that everything is the same. Cool. So let's zoom out then and do just a little bit of like, where does this leave us? I guess going back to the top, you know, you, we were kind of talked a little bit at the beginning of around like, what do we want from language models? And you know, today, man, they're getting awfully good. But we do still have a bunch of, you know, I certainly have like learned a bunch of habits over time for how to use them, where I kind of implicitly am building my practices around some of their limitations here so as to play to their strengths and, you know, not get stuck in their weaknesses. How do you as this paradigm begins to mature and we get like more continual learning, what do you think the experience starts to look like when it comes to things like, what does it mean to start a new chat, you know, and should what sort of relationship do you think people will have with these systems? I, I can imagine like people might have really long running, you know, relationships that when we talk about LLM psychosis now, like that could get really strange. And you know, even more, I mean, the, the, the relationship could be even more compelling. And, you know, the problem could in some ways could be exacerbated by the fact that the models are better. On the flip side, like sometimes I might still want to start fresh, right? Because I might think, jeez, you know, all the stuff I've done with this model in this One Direction, like probably isn't going to help me over here. So maybe I do want to start over in some cases. And then there's just like the question of model upgrade cycles themselves and sort of how do we run evaluations today? We have like Anthropic putting out hundred page reports on every new major model release. And so that paradigm of like, you know, we're going to really take our time to understand these artifacts as as deeply as we can, which I and I wish other companies like DeepMind doing a pretty good job of that open eye doing a pretty good job of that. Some other leading developers not doing much of it at all. I see a lot of virtue in doing all that work, but then I try to port that onto this paradigm and I'm like, well, jeez, you can't like run your full eval suite every single time stamp. So like, how do you think about like what constitutes a version? And you know, when would I change a version? It seems like the a lot of the sort of rhythms of both use and like versioning and deployment and releases like all these things could really be complicated in in a paradigm of like really powerful continual learning. So how do you imagine some of that stuff shaking out I. Think one simple case is that definitely the model gets gets better and better and better in, in understanding what user wants and also adopting themselves to their like their style. And you know, for example, this person when asked about one specific concept, might not expect the same thing when the another person asks the same question. So the model needs to really understand how, how it needs to answer one specific question, but for different people. I, I think that's just definitely gets better and better. If, if we could come up with the continual learner. And on the other hand, you know, we have seen that when we can increase the context window of the model, they, their performance, everything that we know ranging from, you know, all the coding tasks or for example, all the, you know, mathematical reasoning, generally reasoning task, or for example, all of the benchmarks that usually are used for evaluation of the model, all of them gets much better in all those benchmarks. And continual learning can can somehow be seen as as a form of like enhancing the long context understanding of the model. Why in why in the, you know, should like emphasize that the concept of long context understanding or or generally long context of the term of long context is very different from term of continual learning. But continual learning is a super class of lung context. And so potentially if we could come up with a continual learning, then it also has more ability long context understanding and potentially better performance in all of the benchmarks and evaluation that's we are aware of today. So, yeah, yeah, I think that's just one thing that I expect from from continual learners do.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.