Evidence receipt / recommendation
Published · transcript-backedAndrew Gordon Wilson: recommendation
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I think here's here's the way I'll put the question is and I ran into this as a practitioner, you know, back in the day. So maybe I was just stuck in the hump of having, like, too many but not enough parameters, you know, to to get to the double descent phase. I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all. Right? Like, I mean, like, we kind of, you know, we'll talk about marginalization and kind of Bayesian Bayesian perspective on that. So, like, overfitting can be a real problem. So for example, you brought up, you know, conservation. It's like, well, if I'm doing a model and I don't enforce conservation of energy. And then as a result, I end up with some small parameters that cause, like, a little bit of feedback and increasing, you know, energy every single time a robot, you know, takes some action that can cause it to, like, spin out of control. Right? Where I actually did need it to conserve energy and not have that that positive feedback. So I guess I guess maybe we're still struggling with, you know overfitting can be a problem. It's a real problem. It's a known problem. How do we know if we're overfitting in a bad way? We maybe haven't don't have enough parameters. We're stuck in kind of the the area before we got to double descent. Or, like, how do you, in practice, avoid the actual consequences of bad overfitting? Right. So overfitting is real, absolutely. But the conventional wisdom about how we should approach it, I think, is fundamentally misguided. So and this is rooted in things like the bias variance trade off. Like, let's constrain our hypothesis space so that we can't have a bad fit to the data that will make bad predictions and so on. Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit. And interestingly, the prescription is almost the opposite of what people think it it perhaps should be in principle. Like, build a smaller model is usually the the prince the sort of the the prescription for avoiding overfitting. Or to enforce of simplicity.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.