Evidence receipt / evaluation
Published · transcript-backedYi Ma: evaluation
13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]
“You never over so compression, by nature, if the operator are performing compression or denoising, which means this process will no longer overfit anything, right, if you conduct it right, if you converge, the solution will converge on the structure”
— Yi Ma
Source trail
Everything needed to verify it.
- Speaker
- Yi Ma
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Right. I love that characterization. It's similar to the least action principle in in physics. Exactly. In in a sense, we solve problems by taking many, many steps in different directions. I I think we still leave a little bit of entropy open. We don't do pure hill climbing. Yeah. But collectively, we acquire these stepping stones, the totality of that process is we solve very complex problems. But I wanted to touch on you raised a very interesting point, which is we noticed that when we have very large deep learning models, they tend to almost self regularize, and they they learn better. And there's this phenomenon of double descent and all of this. Tell me about that. Fascinating question. Actually, this question actually needs to really bring me back to early days when I tried to understand deep learning. When deep learning arises, there's a lot of phenomena we try to understand. I'm 1 of those, right? Try to understand those phenomenas. There's something good about dropout, something about thresholding, different thresholding, there's something about normalization. Then it also comes to, somehow the model are very big and the parameters are lot. Somehow the deep networks do not have a tendency to overfit. Somehow they still generalize Okay, right? And then, of course, people realize that there's a sort of a rather unlike the traditional classical bias, a virus bias trade off, but there's a double descent. I actually wrote a couple papers about it, and about the normalization, about everything. Around 2000 or late 20 19, I really told my students we should stop, not to explain those isolated phenomena. We only see where like the blind men to elephants. Each 1 say a little piece. Each theory try to explain a little bit. I think there should be a total explanation to this if we get the big picture. All this are just the consequences or implications of that. Suddenly, you know, at that time, we start to touch upon concept of maybe that the process of deep network optimizes something. The layer wise is realizing optimization. They are optimizing objective that promoting parsimony, promoting no dimensionality. Once we realized, actually, I was quite thrilled. So, then I told my student, from now on, we will no longer write any papers about overfitting. Why? Because if the neural networks is trying to the operators try to compress, try to realize certain contracting map, compress volume. They will never overfit, right? Even I overparameterize, it will never overfit. A simple example, if I have data lies on a straight line, a 1 dimensional curve, whatever, I can embed this 1 dimensional line in a 2 dimension, 3 dimension, or 1000000 dimension. But if my operator is always layer wise, at each iteration, my operator is always just shrinking my solution towards the line, right, in all directions. I would never know of it. Even if I overparameterize the embedded line into billions of dimension, I have billions of parameters, but collectively, all those billions of parameter are all shrinking by solution, pushing the solution, denoise it, compress it towards the line, right? Like a power iteration, just like a PCA, right? Power iteration is regardless in of what the dimension embedded, computing the first singular values, right? It's always powerful, but it's converging with the same speed. You never over so compression, by nature, if the operator are performing compression or denoising, which means this process will no longer overfit anything, right, if you conduct it right, if you converge, the solution will converge on the structure you desire for. That raises a natural question. We were interviewing Andrew Wilson from NYU, and he's got this, know, several papers about implicit biases where you just kind of have a combination of, you know, hard biases of symmetries and and everything in in between. And if what you're saying is true, then why do we need inductive biases at all? Could we not pare back a little bit and just have really big models? No. I see. So this is the thing. Right? Exactly.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.