High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon Wilson: evaluation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“Because that means there are gonna be many more different settings of parameters that are consistent with what we observe, and we're just betting everything on 1 of them.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
evaluation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…So Occam's razor has been mentioned multiple times in in our in our in in the last few segments. I wanna dive into that bit because I love Occam's razor. And I first came to really understand it when I became a Bayesian, and it's nice to have a fellow, you know, Arch Bayesian and talk to talk with here. Because in Bayesian inference, if you will, Occam's razor has a very explicit form, like mathematical form, and it's marginalization. Right? It's saying, like, if you have all these parameters around and you you compute the average, you do this integration over all these parameters, that really that's telling you, like, the probability of the model, like, your, you know, full parameter space having been integrated away. And that's really where you get, you know, let's say a penalty for complexity comes into play with with marginalization. It's like if you make if you expand your model's flexibility or you make it more complicated without any gain in inference, well then that kind of, like, counts against you or without any gains of some weird simplicity in the form of, you know, less curvature, you know, things like that. And I was looking at some of your talks, which were really great from from 5 years ago, these these Bayesian tutorials, Bayesian deep learning tutorials you had, I think, when you first when you first went to NYU. I really enjoyed them. And there was a lot of talk about marginalization and the importance of it and the intuitions that you gained from that. And I think but more recently, that's played less of a role in the conversation or people have just given up on trying to do the integrals, and they're just back to doing maximum likelihood and maybe with some hacks and the objective function. So I'm just kinda curious, you know, as a Bayesian, your journey going from understanding, like, the beauty of marginalization and the Occam's razor built into Bayesian inference versus what you do as a practitioner, what you see people doing in practice, how it's playing out in the field or not? And maybe is there a future in which marginalization and doing these computational intractable integrals or something could play a role in driving, like, even better, you know, machine learning and inference? I'm so glad you asked. So there's almost nothing I like more than talking about Bayesian inference. And I'm so glad you said marginalization because I feel like that's often overlooked. Like when someone says Bayesian, probably the word that comes into your most people's mind is prior. And is the prior good? And how do we know what would a good prior be, etcetera. And they get worried about that. But really, the prior is not the defining feature of what it means to be Bayesian. It instead, being Bayesian means that you want to represent the honest belief that you have uncertainty over what solution is correct given a finite data sample. And so if I'm trying to estimate, you know, the bias of a coin that I'm flipping, if I flip it once or twice, regardless of whether it comes up, you know, heads twice or something like this, I ought to have some uncertainty over the bias still. And that uncertainty is manifested through this procedure called marginalization. So I think we can think of regression to get some intuition of this. So we could imagine like a bunch of different points on a sort of y x plot. And maybe the points look roughly like they're on a straight line, but not quite. We could also imagine that there are many different curves that will perfectly run through all of those points, that all look different from each other. Some of them will be very wiggly, some of them will be slowly moving, etcetera. Some of them will basically just be like a straight line. And we wouldn't be able to say, well, we know with 100% certainty it's this curve that's the right description of our problem. But that's exactly what we're doing, you know, 99 a 100 minus epsilon percent of the time when we're training models in deep learning. That is not an honest representation of our beliefs, and it's gonna become a bigger and bigger problem the more expressive our model actually is. Because that means there are gonna be many more different settings of parameters that are consistent with what we observe, and we're just betting everything on 1 of them. And probability theory says, no, that's that's just wrong. That's not what you should be doing. The summoned product rules of probability say you should be doing marginalization. And so that's all to say that Bayesian marginalization, basically looking at all possible solutions that can be expressed by your model class, weighted by their posterior probabilities, is going to be most important when we have a model that's very expressive, has a lot of parameters, so deep learning for example, especially relative to the number of data points that we're considering. And I think often people think of it in in the opposite way that like, oh, maybe Bayesian methods are most relevant in the realm of classical statistics. Like if you're doing some logistic regression and and you know, maybe you want to represent some uncertainty etcetera, but it's not really really suited for deep learning. It's really the opposite. s. Like if you're doing some logistic regression and and you know, maybe you want to represent some uncertainty etcetera, but it's not really really suited for deep learning. It's really the opposite. The challenge then is how do we do this in a way that's computationally tractable. And I think, like with most things in life, the answer of course is nuance. So it's not that like we can either be fully Bayesian or not Bayesian at all. Let's just try to do what makes sense given the resources available to us. And so, like, if we're just doing standard training, you can actually view that as a very crude form of marginalization where you're saying, okay, I'm representing the the posterior as a point mass around the most likely setting of parameters. And then we can say, okay, well, maybe the posterior, which is really like the loss functions that we're minimizing are basically negative log posteriors. Maybe the posterior looks nothing like a point mass and it looks nothing like a Gaussian. It's very multimodal. It's very messy. But we can still do a better job of representing it with a Gaussian than with a point mass. And so let's use a Gaussian approximation. And indeed, when we do that, we often see better generalization because we're representing all these other complementary and compelling explanations to our problem. And then we can keep taking it from there. Okay. Maybe we can develop some MCMC procedure that will explore the loss landscape and capture something much richer than just unimodal Gaussian structure, etcetera. And again, we see improvements in performance in doing that. And so I actually think that Bayesian methods have been an extraordinary success story in deep learning and beyond. But we don't hear as much about them now as we did maybe 10 years ago. And I think there are a number of reasons for that. So to some extent, they're a victim of their own success. So I think there there was some low hanging fruit in being able to do better approximate marginalization in neural nets around 2015. And there was really a lot of progress between about 2015 and 2020 in achieving increasingly better results that will also be kind of computationally efficient. So there are procedures like 1 we developed called SWAG for instance, which was based on some insights we had into the the structure of these loss landscapes that will allow you to do Bayesian marginalization without really any significant additional cost on training. It's a bit more expensive test time and so on. It's also deep kernel learning, which basically just requires a single forward pass through the model, and you get sort of some representation of a specific certainty. And these procedures were adopted in practice. And so if I go like to some more domain specific conference, like some workshop or conference on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence