High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

“I think I want to spend one more beat on what you mean when you say generating its own value, because I'm kind of like, OK, I, I know the transformer architecture pretty well.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
3 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…And so that's Hope architecture, it's self modifying Titan plus continue memory system. I think I want to spend one more beat on what you mean when you say generating its own value, because I'm kind of like, OK, I, I know the transformer architecture pretty well. We've got these like KQ and V vectors, right? And the training process modifies all of those over time. And so the, you know, the sort of general heuristic that I have is like for each token, there's the query vector that sort of indicates like what this token is looking for. There's the key vector that sort of helps indicate like what what other tokens have to offer and, and sort of aligns those right and finds where there's relevant, where there's a match, basically where there's relevance. And then the third one, the value sort of brings up the concepts that then get fed on into the downstream layers that, you know, make sure we have like the right activations for continued processing from there. But all of those are are learned, right? All KQ and V are learned. So I'm not entirely clear on what you mean when you say that the model learns its own values because like doesn't the transformer sort of learn its own value vector as well? I'm not quite clear on the distinction that you're making there between what you know. Well, let's assume people are at least generally familiar with the transformer and know how that goes. So generally in transformer or more accurately softmax attention, what is happening there is that we have a projection of QKV and then the output goes to attention. So attention doesn't have any control on the QKV projections. But when I'm saying that the model or or generally the associative memory tries to generate its own value and then map keys to the value, I mean something like gradient descent. So if it if we just recall the gradient descent, we can see that we have something like the previous state of the West, which is WT minus the gradient of the lost function that we have. So this is equal to the next state of the weights that we have. So if we look at this process, you can break the gradients using the chain rule and write it into the gradients with respect to the output times the inputs data. And now you can see that this is getting a form of associative memory very similar to linear attention because it's WT plus one equals to previous state of the West minus K&K here is XT and then V which is the gradient with respect to the outputs. So if you look at this process, it's very similar to linear attention, but the OR any linear recurrence model. But the very interesting part is it is different from linear attention because the point is if you look at the value, which is the gradient with respect to the output, this gradient with respect to the output is a function of WT, is a function of the current state of the weights. So basically the value, the you know keys and values that we have in associative memory, the value component is not from another component before this recurrent formula, but it is generating by this recurrence recurrent process every time. So that's what's what is going on with this like self refresher process. So in memory, let's say that in in a very simple version of like self modifying Titan. If we want to design this as a simple Titan, what is happening is that I have this X and then QKV projection. I project it into QKV and then pass all of them to the Titan module O that's a simple Titan module. But if I want to have a self modifying Titan, then all of these parameters of the projection of QK and V are optimized inside the Titan maju. So basically the model has the control of somehow modifying its own updating room and somehow generate its own value for for the memory. That's the main difference of these two.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence