Evidence receipt / observation
Published · transcript-backedYi Ma: observation
13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]
“Because there is still compression, then you get a convolution, naturally, as the structure for the compression operator.”
— Yi Ma
Source trail
Everything needed to verify it.
- Speaker
- Yi Ma
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 13 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…you desire for. That raises a natural question. We were interviewing Andrew Wilson from NYU, and he's got this, know, several papers about implicit biases where you just kind of have a combination of, you know, hard biases of symmetries and and everything in in between. And if what you're saying is true, then why do we need inductive biases at all? Could we not pare back a little bit and just have really big models? No. I see. So this is the thing. Right? Exactly. So this is the thing. I was you know, early on, people don't understand deep networks, and there's a lot of empirical trial and error. People tend to use the phrase inductive bias to, either as some kind of magical sauce that either explains the failures or success of you do a certain way to the neural network's design, or how do you train the neural networks. To be honest, for a long time, I never understand what the inductive bias is. And maybe some recognizations, some people reciting some structures about network, about the data. But nowadays, in my recent work, I said that probably, at least from what I understand, all the inductive bias should be formulated as first principle, right? At least from we were able to, for example, deduce all the different network architectures, including the recent white box crate or transformer like, or Red Dot like, ResNet like architecture, or mixture of expert like architecture, all from the only inductive bias is assuming your data distribution you are pursuing are low dimensional. Okay? You can already get the form, the main architecture, or form of operator for each layer. As a REST structure, mixture of XPRO structure, and those operators per layer are precisely conducting denoising, compression, or contrasting. Are there additional assumptions you can make? Yes, you can. For example, if my job is not just to compress the data as it is, I also wanted to induce I wanted to, for example, in object recognition, I also want to enforce, make all data. I wanted my classification to be translational environment, which is symmetry, right? If you allow my task will be environment to certain group action, I want to compress them together. Voila, what do you get? Because there is still compression, then you get a convolution, naturally, as the structure for the compression operator. So convolution is not what we impose upon. It actually results from the first principle, the quote unquote inductive bias assume. You want to compress your data. Also, you want your compression to respect translation environment or rotation environments. That's the result. That is the characteristic of the compression operator for you to achieve that task, right? So there's a lot of so we don't want to build in the inductive bias while we're searching for the solution. The inductive bias, in my understanding, should be the very assumption we make in the very beginning. The rest should be detection. The rest should have no induction anymore. Otherwise, we're doing trial error, right? The inductive. So basically, we build a theory, we should have done all the inductive observations, experiment, and assumptions already. The good theory should start with the very few inductive bias or assumptions or axioms, then the rest should be deductive. I call that first principle. We've been speaking about parsimony, which is what to learn, and self consistency is about how to learn. And we can sketch out a journey, I suppose, from control theory to learning. And and also this this methodology has some interesting, results, I think, around, you know, the continual learning problem. So let's sketch that out. You can see. Right? So the the compression…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.