Evidence receipt / evaluation
Published · transcript-backedYi Ma: evaluation
13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]
“Though, from our lesson, we realized, indeed, actually, that actually those option function has very benign landscapes.”
— Yi Ma
Source trail
Everything needed to verify it.
- Speaker
- Yi Ma
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…You know, when your data lies on very low dimensional sparse surfaces, planes, right, low dimensional planes, orthogonal planes, or load rank matrix, right? And in there, learned a very big lesson. The object function to evaluate those sparsity or low dimensionality, those functions are highly nonlinear, non convex. But yet, you know, traditionally, in all our orthodox understanding about the non convex optimization is they're always hard, right? And in the general classes, NP hard, and there's lots of local spurious local minima. You get stuck with local minima. You get there are some stagnant critical points, flat surface. So basically, the worst picture is very worse, right? It's a nightmare. But through the study of those low dimensional structure, sparse structure, that's what's actually featured in my previous book, right, high dimensional a low dimensional structure of high dimensional data analysis, we actually realized that if a lot of long convex problem, even the optimization problem had long convex landscape, If those problems or even those measures arise from nature, very natural resource, those structures actually are very highly regular, highly has symmetry. The landscape actually are extremely benign, right? Quite contrary to our common understanding about a nonlinear optimization at all, right? This is a complete 180 degree flip of views. In fact, even the higher dimension helps. The higher the dimension, the better. We call it a blessing of dimensionality. So, those regularity, those symmetry will tell us the landscape of this object function are actually beautiful, right? And first of all, they're highly regular. There's no stagnant. There's no flat surface. There's no too many spirits, local minima. And even the local minima, they already have very clear geometrical statistical meaning. And hence, those landscapes are very amenable for very simple algorithm to find the optimal solution, such as a gradient descent, which almost indirectly explain why even we're doing even modern training neural networks, and many more, we're searching low dimensional distribution in very high dimensional spaces. But somehow, gradient descent always end up with somewhere nice. Okay, yeah, fine. You can run a long time, but somehow you always end up with those landscapes are not that hard to traverse, right? So, it actually could be precisely because those object function are highly regular. Hence, now get back to the read reduction object function, right? If you look at the object function, it's not something arbitrary, right? It's counting the volume of the whole minus the parts, right? It's something extremely objective, right? It's not like a loss function people come up with randomly, oh, add this term, weighted sum, add different weights. You know, if you use this, you know, sort of an empirical penalty or empirical or even some kind of ad hoc. So, all the terms are describing physical volumes of the data, d sum, add different weights. You know, if you use this, you know, sort of an empirical penalty or empirical or even some kind of ad hoc. So, all the terms are describing physical volumes of the data, right? Hence, you should expect those are the quantity arising in nature. Though, from our lesson, we realized, indeed, actually, that actually those option function has very benign landscapes. Even the local minima, not only the global minima, corresponding solutions that are give you orthogonal subspaces, even the local ones, right, they have lot of global optimal ones. They have similar geometric structures, right? And there's no other weird critical points that will slow down the search for those minimas. So, that's actually quite interesting. So, you can see this revelation allow us to understand, right? Where maybe intelligence is precisely explored and harnessing those things. So there is actually a misunderstanding about, you know, last 2 years. When we understand intelligence more and more, there is a very big misunderstanding about intelligence, right? In machine you study, you know, machine learning theory, right? We have a tendency to believe intelligence, especially intelligence in nature, is designed to solve the hardest problem, the worst case. I actually beg to differ. Intelligence is precisely the ability to identify what is easy to address first. What is easy to learn. What is natural to learn first? The only 1 that has been done and the resource permit is start to get into more and more advanced tasks. Not everybody needs to learn advanced mathematics to supply. Animals don't, right? Nature find what is the easiest things with minimal energy, minimal effort almost, to learn the most knowledge so they survive the best, right? Again, this is the principle of parsimony at play. There's another level of resource parsimony at play here, right? Hence, once you realize this, you're realizing understanding intelligence should be really understanding what's really most common, right? The low dimensional structure, most easy ones, smooth ones, benign distribution, easy to get, get away with a few samples, fewer samples, right? And very easy to formulate. In fact, that's how science progressed. A lot of the physical models, Newton's law, they're very simple. The discover simple ones. Then we gradually reach through general relativity and then to quantum mechanics. Those equations get far more and more complicated later, right? So this is hence, it's the same process. We identify what is the most common first, right? What is the most easiest task first? Hence, we don't want to many of the machine learning theory tends to derive a bond for the worst cases. I think that's we probably should think twice. Right. I love that characterization. It's similar to the least action principle in in physics. Exactly. In in a sense, we solve problems by taking many, many steps in different directions. I I think we still leave a little bit of entropy open. We don't do pure hill climbing. Yeah. But collectively, we acquire these stepping stones, the totality of that process is we solve very complex problems. But I wanted to touch on you raised a very interesting point, which is we noticed that when we have very large deep learning models, they tend to almost self regularize, and they they learn better. And there's this phenomenon of double descent and all of this. Tell me about that. Fascinating question. Actually, this…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.