Evidence receipt / recommendation
Published · transcript-backedAndrew Gordon Wilson: recommendation
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“There are very subtle issues, think, with some of the ways that this can be done, and so we wrote a paper all about that. But largely speaking, it's, you know, something that I think people should acquaint themselves with because it it really is sort of getting at something very fundamental and it's has extraordinary practical value.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Exactly. Yeah. Maybe flatness in random directions versus sharpest directions. Maybe it needs to be parameterization invariance. Like that's that's been often a criticism of of discussions around flatness etcetera that it's most ways of measuring flatness aren't parameterization invariant. I have thoughts about that separately, but anyway, you you might then look at Fisher matrices or something. Okay. So then once you've decided on what it means to be flat and no 1 will completely agree with you or most people will will strongly disagree whatever you choose, then you have to decide like how much am I gonna penalize sharpness. No 1 will really know what to do about that either. Right. And like the idea is you wanna build a loss function that is a better proxy for generalization by accounting for things like flatness, but it's just this really messy rabbit hole. Whereas if you're just doing marginalization, this is all happening under the hood. You don't need to worry about it. And so that's really elegant. And then there's Occam's razor in terms of model selection. And I I'd recommend that everyone check out chapter 28 of David Mackay's information theory and and inference learning algorithms book. It's titled Occam's razor. And I don't agree with some of the stuff in that chapter and in fact we wrote a paper about this on Bayesian model selection and marginal likelihoods. But it's it's a very beautiful description of automatic Occam's Razor and he has these just extraordinary visualizations to try to demonstrate how that's possible. So he has like these also, you know, everyday examples as well like this idea that maybe you have like a block behind a tree. But if you had x-ray vision, it would appear to be like 2 blocks of equal height and color or 10 blocks or something like this. But given that you don't have x-ray vision and you don't think this is a trick question, you'd be pretty confident it's gotta be 1 block. And this seems like some manifestation of Occam's razor and you might try to rationalize this by saying, well, it would be just a remarkable coincidence to have 2 blocks standing next to each other of equal height and color. But he argues that this is actually a quantifiable consequence of doing Bayesian marginalization. And so he basically has this conceptualization where you have all possible data sets on a horizontal axis, and the probability of generating a particular data set under your model on the vertical axis. And that's the marginal likelihood, the the probability that you would generate your training data under your model prior. And so if you have the 1 block model, you're not gonna be able to generate very many different data sets. So most of the datasets on the horizontal axis have no support, no p d given m. But the ones that it can generate, it's gonna have to give a lot of probability for it because this is a proper, normalizable probability density. the horizontal axis have no support, no p d given m. But the ones that it can generate, it's gonna have to give a lot of probability for it because this is a proper, normalizable probability density. Similarly, if you have the 10 block model, you can generate all sorts of different types of observations, you're but gonna have to sort of spread that mass more thinly because this is like a proper normalizable probability density. And so for a particular dataset that's consistent with both of these models, like the block behind the tree example, the 1 block model is actually gonna have significantly more probability. And this is even not factoring into account the idea that maybe we have some prior preference for simplicity and that there's like, we think that 1 block is more likely than 10 block or something like that. He's just let's just forget about this ratio of prior odds and only consider this ratio of marginal likelihoods. And so it's just a beautiful demonstration of how Bayesian inference automatically encapsulates a notion of Occam's razor. This is just a fundamental question in science. Like, if you have arbitrarily many hypotheses that are consistent with any number of observations you can ever record, how do you choose between them? What what what's the principal way of approaching that problem? And the answer to some extent, I think, is given by Bayesian marginal likelihoods. There are very subtle issues, think, with some of the ways that this can be done, and so we wrote a paper all about that. But largely speaking, it's, you know, something that I think people should acquaint themselves with because it it really is sort of getting at something very fundamental and it's has extraordinary practical value. What was that paper? Was that…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.