Evidence receipt / preference
Published · transcript-backedAli Behrouz: preference
3 Jun 2026 The Cognitive Revolution Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
“When we are talking about in context recall tasks or or generally like recall intensive tasks, in my opinion all those tasks are designed for Transformers.”
Source trail
Everything needed to verify it.
- Speaker
- Ali Behrouz
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 3 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, got you. I'm always a big fan of trying to get a little bit better sense for the micro skills of different architectures. So for example, of course, you know, Transformers, because the full sequence, the full context is in working memory at all times. It's pretty hard to beat. And I, I feel like you even sort of have like kind of a theoretical argument now that like it maybe even be kind of impossible to beat in some, in some of these tasks where the idea is like recall from the context window. But then, you know, we saw things with Mamba, for example, where it was better at learning from like sparse signal. And this was sort of a micro skill that it, that architecture excelled at that the transformer relatively struggled with. What, what, what have you seen in the hope case? You know, are there little micro? And I, I think it's, it's very interesting because it does kind of ladder up to the overall performance, you know, and what these things are actually good or bad at, right? I mean, the ability to recall something in, in context is really important when you need it. The ability to learn from or kind of filter out noise and, and get to the, you know, the signal that really matters is, is really important when you need it. So are there particular micro skills that that stand out to me that the language translation one is, is an interesting one in a macro sense of like that's a hard task. But I wonder if you drill down to these like very micro building block competencies that models or architectures can either have or not have, what stands out in terms of what this has that Transformers don't have or don't have as strongly? When we are talking about in context recall tasks or or generally like recall intensive tasks, in my opinion all those tasks are designed for Transformers. They are not designed to compare architectures, but they are specifically designed for Transformers. Why I'm saying that? Because you cannot expect from a model or even a human to perform needle in haystack perfectly or for example do some recall intensive task. For example, I assume that you have a code and like couple of 1000 lines of code and then simply you want to recall what was the value of X at some line of the code. And so it's almost impossible or or at least it's very, very hard for a human or even like for other models to do that. And but on the other hand, it's it's pretty much simple for Transformers because they have direct access to the entire history in their context. And so it's very simple to just like, you know, find that token and pass it as the out and in fact find it somehow in recall intensive task like this in context recall task that we have here. Actually the gap between recurrent architectures which actually they perform as is also very great. If you compare the, you know, first generation of recurrent architectures to the transformer, we can see that these gaps was much, much larger. Now this gap is getting like smaller and the performance of other recurrent models is also really great. But the interesting part for me was that Hope at least closed this performance, closed this gap in the performance of the model compared to Transformers. While they are not expected to do that. Transformers is maybe expect Transformer to do that because it has a tension block, but we don't expect a compression based model to perform recall task. And so I think somehow it was interesting for me. And what does the mad data set get at? How should we understand like because that's, that's where to just replay back to what you just said. The on these like needle in a haystack. These like very difficult recall tasks from earlier in context, the transformer remains the best. The recurrent models, which only have some latent representation and and don't have the ability to look back at the original raw text don't perform as well. But with each generation of improvement, and here you've got several, the hope architecture does the best of of the recurrent ones that don't have the the full explicit context in working memory at runtime. And then so that gap is closing, flipping over to the mad data set. Here the hope architecture is performing better than everything including the transformer. What like micro skills, is that testing which? Which should we take away from that result? The.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.