Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“I mean, maybe it's been somewhat fruitful, but I think also people would get very confused and they try to interpret dreams or understand, you know, what's going on there.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…From the technical point of view, we cannot like grow the model to like arbitrary large number of parameters. But the point here is that it's like a periodic process. We add some parameters and then we free them for the next step of consolidation. When we are in the first log, we add some components. When it reaches its capacity, it means that it is the time that we need to consolidate the memory to the next step. And when we consolidate all this knowledge to the next step, we just remove all the extra capacity that we have added to this level and freedom for the like other levels like faster levels, so they can also consolidate their their memory to this block as well. So generally it's like a periodic process. We add components and remove them add. Components and I see. Gotcha. OK, interesting. And what more can you tell us about the dreaming phase in terms of like just a little bit more practically what's going on there? I mean, it's certainly when I try to introspect into dreams, it, I think it hasn't been super. I mean, maybe it's been somewhat fruitful, but I think also people would get very confused and they try to interpret dreams or understand, you know, what's going on there. So I won't even attempt to, you know, ground my understanding in my human dreams, which seemed like quite a, quite a hard thing to untangle, I guess. But here there's a more. I mean, you got to design the process. So like what procedurally, what is going on in in dreaming just a little bit more like mechanically, procedurally? The concept of dreaming doesn't mean that it's exactly the same thing as dreaming in human. It's just, you know, at a very high level they seems to be very similar. And so in that's that's one point. And another point is that again, the concept of sleep and dreaming for a language model might be very different from the concept of dreaming and asleep for the, for example, vision model. Because potentially a vision model will generate some images generative like a vision model might might generate some images during dreaming. While in this case of language modeling we are generating text. But the framework is very general. It can adopt to any, any data modality. And so that's that's very general. But the point is here when we are like doing that for language modeling, what is happening there is we generate some context, we generate some text. And how do we generate those texts? It's it's on policy, the solution the same way that that we discussed earlier, we have a model, we copy that, you know, we want to distill the knowledge from one level to the other one. And so we free the parameters of the slower level and so on, so forth. And so we ask the smaller one, the small model which has the knowledge of the context as well into its parameters. We ask that to generate some text and then we want to train or or somehow you know, update. It's, it's a better chance to use update the actual model, prompt the actual model parameters for on these specific data set that is generated on by the model. And then how do we train it? We start with one part of the sequence, just sample some some of the tokens and then ask the model to predict the next tokens in that sequence. O that, for example, that's really similar to, you know, generating some, some synthetic data. It seems if the model can perfectly predict the future tokens, it means that it has already knew about the knowledge that is stored in the previous block. So it's it's a perfect model. But if it cannot properly, you know, predict the continuation of the sequence, it means that it doesn't have the knowledge that is stored in the context and needs to update itself to understand that knowledge as well. So it's it's somehow a form of on policy distillation that is happening inside the model. And so it, as I mentioned again, just just as a summary, we have two phases. 1 is the generation which generates some text about the knowledge in the context. And then the second part is on policy distillation that we distill the knowledge from one level to the other. That's what's happening in the dreaming phase. I mean we also have the self modifying part as well, but I think that's that's the main idea of memory consolidation. And how? How did dreaming happen here?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.