Evidence receipt / prediction
Published · transcript-backedYi Ma: prediction
13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]
“If you think about the whole diffusion denoising model, right, people are very popular right now to do why do we add noise to data, right, and to the whole world? Because we don't know where the distribution is, right?”
— Yi Ma
Source trail
Everything needed to verify it.
- Speaker
- Yi Ma
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 13 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…should introduce the coding rate formula. Did have a question about that, which is there is a, there's an Epsilon on there. So there's a bit of a question of how how do we tune that and and what does that mean. We should also bring we've been talking about this a little bit, this this concept of an LDR, so a linear discriminative representation. And just more broadly, with these inductive priors, there's always the question of when we do abstraction to model regularities in the universe, there's always a little bit left over, isn't there? So to to what extent can we think of these things as natural? Actually, you touched upon about the Eepstow. You touched upon a very, very deep question. We actually It actually took me almost 30 years to understand it, to be honest, right? We did mention that early on when we do try to differentiate different, measure different volumes. It turns out lossy coding is necessary. It's not just something that is something hacky. It actually turns out to be necessary to do lossy coding. In fact, recently we started to realize that noise actually plays very different roles. And yet, it's very confounded, very confusing to many people. Even this is something actually my students will sit and realize. We actually probably will have some papers about it. I can elucidate this a little bit. If you think about the whole diffusion denoising model, right, people are very popular right now to do why do we add noise to data, right, and to the whole world? Because we don't know where the distribution is, right? So, there is a phrase everybody knows, all roads to Rome, right? So why is that? Has anyone given a thought to why all roads to Rome? Because, very simple, at some point in history, Rome builds the road to reach the whole world, right? That's a diffusion process. Then if you want to know Rome, then you do the denoise. It follows the same way back. You get to where the Rome is. So that's the node dimensional structure. That's where the knowledge is, right? That's the OSS is, right? So hence, it's a very natural process that we add noise. It's adding noise to the precisely building the roads. And the denoising brings us back, remember where we come from, and so on. And that's a big epsilon. We have to add noise to reach the whole earth. There's another actually, there's another noise, right? Remember, we only have isolated sample. Even we talk about manifolds, right? But how many points do you have on the manifolds? How many points do you observe? They're always finite, right? But why do you call it a continuum? Why do you collect dots as lines, planes, surfaces? When do you do that? Hence, noise plays another role within the manifold. Even you have finite samples, if you allow lossy coding, if you allow packing spheres in that, things start to connect. You start to connect. Noise is very important to help to connect the dots, right? We all know the phenomena of percolation, right? We see raindrops on the floor. You only see 2 phases, right? 1 phase is all the dots are isolated. Another phase is all things get wet. You never see anything in the middle because there's a sharp phase transition. Once the sphere once the dots, the density gets high enough, it collects everything, right? Maybe that's a phase transition we reach we realize a connected plane is a better solution to explain all the data, more parsimonious, more economic. The cost to memorize all the dots versus to memorize the other plane start to switch. Maybe abstraction has something to do with that. I don't know. But from a compression point of view, this can really allow us to explain when do we go from 0 dimension samples r plane start to switch. Maybe abstraction has something to do with that. I don't know. But from a compression point of view, this can really allow us to explain when do we go from 0 dimension samples to prefer a low dimensional manifold. And also, how go from that low dimensional manifold to reach the rest of the world? Right? So you can see, even in this process, noise is already playing epsilon plays different roles. And at some point, they get collected, right, around the surface. That's why we're still trying to figure out what happens. But at this big 2 phases, we already know, right? The role of epsilon actually plays different roles, right? And I think definitely in the past many years, our understanding about the subject, how do we compress, how do we pursue the low dimensional structure from finite samples, it's quite our understanding about this problem has truly advanced dramatically. I'm very happy. Honestly, this is a question baffles me when I was a graduate student. You can see even my early work about the lossy coding, lossy compression reflected my baffledness about it. And I really feel very thrilled that I recently started to understand…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.