Evidence receipt / evaluation
Published · transcript-backedYi Ma: evaluation
13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]
“I think we're sort of managed to do that, at least for the structure we've discovered so far, provide rather unified explanation to what they have done.”
— Yi Ma
Source trail
Everything needed to verify it.
- Speaker
- Yi Ma
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…and you made some very interesting discoveries. So for example, multi head self attention can be derived as a gradient step on rate coding, and also MLPs as sparsifying, specification operators. And and also you were talking about how something like a transformer could be described in a principled way. So there's there's this interesting thing, isn't there, that we we we designed it even well, we we didn't even design them. We we kind of empirically tried of lots of different things, we happened upon the transformer. But something like that can actually come about from a first principles approach. If you look at the past decade or so, evolution of also, it's kind of natural selection process for the big models. Right? From early days, Alexlet, Lollet, Alexlet, VGG, or then ResNet or transformers. By the way, this is just 1 of those for survivors, right, as I said, just like lateral selections, right? Remember, people don't forget. There's a time there is a very popular area called AutoML, right? People tend to do random search for better architectures, right? Somehow, why only a few survived? There must be a reason, right? They must capture certain structures. They must did something right. Now, from our understanding so far, the ResNet actually capturing the fact that each layer should be doing optimization. The ResNet precisely reflect the iterative optimization architecture, right? And MOE precisely capture fact. We're trying to cluster, compress what's similar, and discern or classify what's different or contrast what's dissimilar, right? And you wanted to develop different experts, right? We call them experts. We call them cluster. We call them a group. So be it, right? And the transformer, again, right, has captured what is the correlation. Self attention has precisely compute what is the correlation in the data, covariance in the data, what's correlated? And using that to further sparsify, further classify things, to organize the distributions. They must do something. They're somehow close to something right, right? So also, it's almost like a belief for us, right? If we believe there's something right, then we should be able to derive, create from first principle, have a very clear unified understanding. I think we're sort of managed to do that, at least for the structure we've discovered so far, provide rather unified explanation to what they have done. To be honest, the early even maybe our earliest motives try to explain to understand what we have done. But once we understood it, we realized that we can go much further, right? And realize even the current architectures, there's a lot of room for improvement. Not only we can dramatically simplify them, you can see in the past after the in the past yes, last year and this year, there's a series of work from my group, right? Really just showcase people, right? You can actually once you understand what is done with the principle, you can dramatically simplify. You can even throw away the MLP layer if you only care about the compression. You don't care about the final representation. And or you can make the attention head. Since we know what is optimizing, it's optimizing the rate reduction object function, then we can find what is the equivalent variation form of that object function, which is much easier to optimize. We end up with a we call it a toss, right? The computing the covariance the self attention step is only linear in the dimension, no longer quadratic, like the current tension is doing. Of course, if you look at the literature, there are other people have found try to identify linear complexity, is only linear in the dimension, no longer quadratic, like the current tension is doing. Of course, if you look at the literature, there are other people have found try to identify linear complexity, such as, you know, or I think there's RWVK or something, so empirically. But again, it's through trial and error. But this is so now we derive this in the math in purely mathematical way, because we just find an equivalent variational form of the same object function that have the same global optimal. But it's just much easier to optimize. This is a trick we do all the time, right? All the tricks, see, in the 200 years plus years of developing better optimization algorithm, all those ideas can help us now to design better operator, descent operators, or optimization architectures to improve the design of current architectures. Honestly, we have not really started that far, right? There's many acceleration techniques, preconditioning, conjugate gradient, which explore different landscape. Once we understand the landscape, the type the cost of object function better, there's gazillions of ideas. We can further improve the efficiency. Honestly, we haven't started that far, right? I mean, that's actually got some of my students excited to pursue this, realizing how little we have done from an optimization perspective, how much room there might still be for improvement. Some my students are quite excited. So you can see, you know, even within the last couple of years, we already have 2 or 3 different generation of architectures that in the past, it's almost unthinkable, because the new generation always come from a different group, right? It's like a random process. Whoever gets lucky, maybe discover something works. Try hard enough to get something to work. It's a tantalizing idea though that through this principle of optimization, there could be, you know, a convergent evolution towards…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.