High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Andrew Gordon Wilson: recommendation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
recommendation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…I think here's here's the way I'll put the question is and I ran into this as a practitioner, you know, back in the day. So maybe I was just stuck in the hump of having, like, too many but not enough parameters, you know, to to get to the double descent phase. I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all. Right? Like, I mean, like, we kind of, you know, we'll talk about marginalization and kind of Bayesian Bayesian perspective on that. So, like, overfitting can be a real problem. So for example, you brought up, you know, conservation. It's like, well, if I'm doing a model and I don't enforce conservation of energy. And then as a result, I end up with some small parameters that cause, like, a little bit of feedback and increasing, you know, energy every single time a robot, you know, takes some action that can cause it to, like, spin out of control. Right? Where I actually did need it to conserve energy and not have that that positive feedback. So I guess I guess maybe we're still struggling with, you know overfitting can be a problem. It's a real problem. It's a known problem. How do we know if we're overfitting in a bad way? We maybe haven't don't have enough parameters. We're stuck in kind of the the area before we got to double descent. Or, like, how do you, in practice, avoid the actual consequences of bad overfitting? Right. So overfitting is real, absolutely. But the conventional wisdom about how we should approach it, I think, is fundamentally misguided. So and this is rooted in things like the bias variance trade off. Like, let's constrain our hypothesis space so that we can't have a bad fit to the data that will make bad predictions and so on. Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit. And interestingly, the prescription is almost the opposite of what people think it it perhaps should be in principle. Like, build a smaller model is usually the the prince the sort of the the prescription for avoiding overfitting. Or to enforce of simplicity.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence