High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Tim Scarfe: observation

13 Dec 2025 Machine Learning Street Talk The Mathematical Foundations of Intelligence [Professor Yi Ma]

“We we just find that structure. And that's why when I watched your presentation, I was very intrigued when you said that denoising, iterative denoising is is a form of of compression.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
observation
Recorded
13 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…r plane start to switch. Maybe abstraction has something to do with that. I don't know. But from a compression point of view, this can really allow us to explain when do we go from 0 dimension samples to prefer a low dimensional manifold. And also, how go from that low dimensional manifold to reach the rest of the world? Right? So you can see, even in this process, noise is already playing epsilon plays different roles. And at some point, they get collected, right, around the surface. That's why we're still trying to figure out what happens. But at this big 2 phases, we already know, right? The role of epsilon actually plays different roles, right? And I think definitely in the past many years, our understanding about the subject, how do we compress, how do we pursue the low dimensional structure from finite samples, it's quite our understanding about this problem has truly advanced dramatically. I'm very happy. Honestly, this is a question baffles me when I was a graduate student. You can see even my early work about the lossy coding, lossy compression reflected my baffledness about it. And I really feel very thrilled that I recently started to understand those things in a more unified, more not only theoretical way, but also even algorithmic way. Yeah. It's it's so fascinating that we can look out the window. Uh-huh. And we we ignore so much detail. We don't look at the leaves on the roads. We we just find that structure. And that's why when I watched your presentation, I was very intrigued when you said that denoising, iterative denoising is is a form of of compression. Yep. I wanted to mention your ICML 20 24, so last year, was in Vienna, right, with with Wang. And you found that when you have lost surfaces using using this technique, they are dramatically different. They're very smooth. There's no kind of harsh local minima and so on. What's the intuition for that? In fact, the phenomena our understanding about those phenomena is actually going back to the early days we studied sparsity. You know, when your data lies on very low dimensional sparse surfaces, planes, right, low dimensional planes, orthogonal planes, or load rank matrix, right? And in there, learned a very big lesson. The object function to evaluate those sparsity or low dimensionality, those functions are highly nonlinear, non convex. But yet, you know, traditionally, in all our orthodox understanding about the non convex optimization is they're always hard, right? And in the general classes, NP hard, and there's lots of local spurious local minima. You get stuck with local minima. You get there are some stagnant critical points, flat surface. So basically, the worst picture is very worse, right? It's a nightmare. But through the study of those low dimensional structure, sparse structure, that's what's actually featured in my previous book, right, high dimensional a low dimensional structure of high dimensional data analysis, we actually realized that if a lot of long convex problem, even the optimization problem had long convex landscape, If those problems or even those measures arise from nature, very natural resource, those structures actually are very highly regular, highly has symmetry. The landscape actually are extremely benign, right? Quite contrary to our common understanding about a nonlinear optimization at all, right? This is a complete 180 degree flip of views. In fact, even the higher dimension helps. The higher the dimension, the better. We call it a blessing of dimensionality. So, those regularity, those symmetry will tell us the landscape of this object function are actually beautiful, right? And first of all, they're highly regular. There's no stagnant. There's no flat surface. There's no too many spirits, local minima. And even the local minima, they already have very clear geometrical statistical meaning. And hence, those landscapes are very amenable for very simple algorithm to find the optimal solution, such as a gradient descent, which almost indirectly explain why even we're doing even modern training neural networks, and many more, we're searching low dimensional distribution in very high dimensional spaces. But somehow, gradient descent always end up with somewhere nice. Okay, yeah, fine. You can run a long time, but somehow you always end up with those landscapes are not that hard to traverse, right? So, it actually could be precisely because those object function are highly regular. Hence, now get back to the read reduction object function, right? If you look at the object function, it's not something arbitrary, right? It's counting the volume of the whole minus the parts, right? It's something extremely objective, right? It's not like a loss function people come up with randomly, oh, add this term, weighted sum, add different weights. You know, if you use this, you know, sort of an empirical penalty or empirical or even some kind of ad hoc. So, all the terms are describing physical volumes of the data,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence