High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon Wilson: evaluation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“ence on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. So this is really useful in practice, but it's very hard to kind of go beyond the useful approximations I think we developed without doing a significantly without making sort of a really significant, you know, 10 year kind of style moonshot landing kind of investment in those directions, which I think is worth making.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
evaluation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…s. Like if you're doing some logistic regression and and you know, maybe you want to represent some uncertainty etcetera, but it's not really really suited for deep learning. It's really the opposite. The challenge then is how do we do this in a way that's computationally tractable. And I think, like with most things in life, the answer of course is nuance. So it's not that like we can either be fully Bayesian or not Bayesian at all. Let's just try to do what makes sense given the resources available to us. And so, like, if we're just doing standard training, you can actually view that as a very crude form of marginalization where you're saying, okay, I'm representing the the posterior as a point mass around the most likely setting of parameters. And then we can say, okay, well, maybe the posterior, which is really like the loss functions that we're minimizing are basically negative log posteriors. Maybe the posterior looks nothing like a point mass and it looks nothing like a Gaussian. It's very multimodal. It's very messy. But we can still do a better job of representing it with a Gaussian than with a point mass. And so let's use a Gaussian approximation. And indeed, when we do that, we often see better generalization because we're representing all these other complementary and compelling explanations to our problem. And then we can keep taking it from there. Okay. Maybe we can develop some MCMC procedure that will explore the loss landscape and capture something much richer than just unimodal Gaussian structure, etcetera. And again, we see improvements in performance in doing that. And so I actually think that Bayesian methods have been an extraordinary success story in deep learning and beyond. But we don't hear as much about them now as we did maybe 10 years ago. And I think there are a number of reasons for that. So to some extent, they're a victim of their own success. So I think there there was some low hanging fruit in being able to do better approximate marginalization in neural nets around 2015. And there was really a lot of progress between about 2015 and 2020 in achieving increasingly better results that will also be kind of computationally efficient. So there are procedures like 1 we developed called SWAG for instance, which was based on some insights we had into the the structure of these loss landscapes that will allow you to do Bayesian marginalization without really any significant additional cost on training. It's a bit more expensive test time and so on. It's also deep kernel learning, which basically just requires a single forward pass through the model, and you get sort of some representation of a specific certainty. And these procedures were adopted in practice. And so if I go like to some more domain specific conference, like some workshop or conference on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. ence on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. So this is really useful in practice, but it's very hard to kind of go beyond the useful approximations I think we developed without doing a significantly without making sort of a really significant, you know, 10 year kind of style moonshot landing kind of investment in those directions, which I think is worth making. Secondly, there's the advent of LLMs, and that did sort of change things in practice. So like we went from million parameter models to billion parameter models. And the role of epistemic uncertainty is also a little bit less clear, I think, in some instances when we're working with LLMs. And so this is just a a real technical challenge and I think we're just sort of barely starting to scratch the surface of how to think about it. But the motivation is absolutely there. I think it's really greater with these big models than it is for virtually any other model class. And in this sense, I don't know how anyone could not be a Bayesian. Like, how could anyone say, I don't want to represent epistemic uncertainty? I don't like using jargon. So there's like, there's aleatoric uncertainty, which means irreducible uncertainty often associated. Like, we have given a data set, we have some noise on that data set, there's nothing we can do about it. That's aleatoric uncertainty. Epistemic uncertainty is uncertainty that's reducible with more information. And we believe that uncertainty is there. And there's even an argument that that might be the only uncertainty in the world. Right? Like things that seem intrinsically random like I could roll a dice Right. And it might seem like, okay, there's a 1 6 probability that it'll land on any 1 of the sides. But if I had enough information, like if I could sort of exactly know the strength of the throw, the wind in the air, the friction on the table, and I had the right physical model, I should be able to predict exactly what side it's gonna land on with each throw. So that's just saying, okay, the more information we gather, the less uncertainty we have, even for this process that somehow seems intrinsically random. And physicists, modern physicists now, you know, probably do believe that some aspects of the universe are intrinsically random, radioactive decay and things like this. But many others have not historically, like Einstein was a determinist. And so he would have believed that the only uncertainty there is is epistemic uncertainty. And I think for practical purposes, that's not an unreasonable belief. And so it's absolutely crucial that we try to model it in some way to not do it is just to do something that's mathematically incorrect and could come at a huge cost. There's also a change I think that's happened in this movement to LLMs in the way that people think about data. So it used to be the case that you kind of have a fixed data set and then you throw everything you can at it to get the best possible performance. ment to LLMs in the way that people think about data. So it used to be the case that you kind of have a fixed data set and then you throw everything you can at it to get the best possible performance. Maybe you have some scientific problem and you just really want to do well, and you're limited by computation and other things to some extent. But it's not the case that you have this sort of trade off to manage for a given computational budget between the size of the data and the size of the model. And so that's described by these scaling laws like Chinchilla scaling laws. And that is starting to become the assumption in mainstream machine learning that you have almost like an arbitrary amount of data and you have a fixed computational budget and you want the best possible performance under that computational budget. In that case, you know, rather than trying to be more Bayesian or something like this, it could actually make sense just to use more data.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence