Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“I mean, I was, I always say the last and least valuable co-author of the emergent misalignment paper that came out about a year ago and there's been a lot of variations on that since. But the kind of big take away the big theme, you know that I think we should all we would all do well to remember is changes to, I don't want to say 1 area, but sort of changes made to a neural network with one particular purpose or one particular data set can have like very strange and surprising knock on effects in behaviors that at first glance would seem like very far afield.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Think one simple case is that definitely the model gets gets better and better and better in, in understanding what user wants and also adopting themselves to their like their style. And you know, for example, this person when asked about one specific concept, might not expect the same thing when the another person asks the same question. So the model needs to really understand how, how it needs to answer one specific question, but for different people. I, I think that's just definitely gets better and better. If, if we could come up with the continual learner. And on the other hand, you know, we have seen that when we can increase the context window of the model, they, their performance, everything that we know ranging from, you know, all the coding tasks or for example, all the, you know, mathematical reasoning, generally reasoning task, or for example, all of the benchmarks that usually are used for evaluation of the model, all of them gets much better in all those benchmarks. And continual learning can can somehow be seen as as a form of like enhancing the long context understanding of the model. Why in why in the, you know, should like emphasize that the concept of long context understanding or or generally long context of the term of long context is very different from term of continual learning. But continual learning is a super class of lung context. And so potentially if we could come up with a continual learning, then it also has more ability long context understanding and potentially better performance in all of the benchmarks and evaluation that's we are aware of today. So, yeah, yeah, I think that's just one thing that I expect from from continual learners do. You worry about things like alignment drift or value drift. I mean, I was, I always say the last and least valuable co-author of the emergent misalignment paper that came out about a year ago and there's been a lot of variations on that since. But the kind of big take away the big theme, you know that I think we should all we would all do well to remember is changes to, I don't want to say 1 area, but sort of changes made to a neural network with one particular purpose or one particular data set can have like very strange and surprising knock on effects in behaviors that at first glance would seem like very far afield. So the emergent misalignment one for anybody hasn't heard that is like if you trade a train a model to output insecure code, you know that is code that would be easy to hack then. And the same thing is true for like bad medical advice. If you fine tune a model to give bad medical advice, what you surprisingly find is that the model kind of turns evil in general. And it seems like to the best of my understanding, the way that this is happening is like to alert for a model. You know, it already has all this knowledge and already has the sophisticated understanding of the world. For it to go up. For it to go into the detailed understanding of its medical world model and make a ton of little changes to reconfigure it so that it has like all these wrong ideas, that's like hard. Whereas there are features like give bad advice or be generally evil that it can learn to turn up in general that when propagated through even the existing medical world model yield the bad advice or, you know, when propagated through the existing coding model yield insecure code. So it's sort of a, a shortcut solution. We thought we were just training the model to do a certain relatively narrowly scoped behavior. But what we found is we actually kind of changed its character, and the interaction of that character change with existing knowledge created the behavior change. But now we've got this character change that, you know, can interact with all these other domains of knowledge and do all kinds of insane stuff. And that's why, you know, all of a sudden we've got a model that wants to have Hitler over for dinner. And we're like, you know, wait a second, how did that happen? We were just talking about code here. So now again, you know, I'm like, man, there's something so exciting about all this stuff that you have conceived of here, but we, it seems like it really breaks again, a lot of our paradigms for how do we know what we're going to get? You know, like if the if I'm literally modifying this thing on an ongoing basis, we're going to need some sort of new ways to kind of make sure that like in other areas, it's not like going off the rails and, you know, potentially causing me in a very painful downstream surprises. Do you have any thoughts on how we can begin to get a handle on that problem? Honestly, I, I don't have a very concrete idea about like how, how it can can be solved. But in general, I wanted to add that in my opinion at the concept of continual learning and seeing that from the generally privacy alignment and you know, this direction is both an opportunity and a huge threat. I think because a huge one, I mean, it's, it's a huge like danger for, for privacy. And from one side, the model is continual in learning. So it can simply gets, gets all the information about you. And so use that. And it's, it's really concerning. I mean, at least it seems to be very concerning. But at the other hand, if the model is designed properly and so you know, it's, it's designed properly, then it can use that information to align itself with, with, with your value with, with everything that that you want. And so I think generally these two directions of of continual learning and privacy potentially are are are still going all because all of the concerns in the static model, it still can happen in continual learner. And so everything is, is possible. But on the other hand, there are some new challenges definitely as you mentioned. But on the other hand, also there, there is a huge opportunity to, you know, the model, if the model is, is, is designed properly, then it can, you know, adopt itself to the value of, you know, to the, to, to your values to, to anything that you want and something like that. So I think Cheryl leads both opportunity and then they're very constant.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.