High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Keith Duggar: evaluation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all.”

— Keith Duggar

Source trail

Everything needed to verify it.

Speaker
Keith Duggar
Attribution
Verified speaker
Claim type
evaluation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…So I think the bias variance trade off is an incredible misnomer. There doesn't actually have to be a trade off. So the idea, I guess, behind the classical bias variance trade off is that your generalization error can be compartmentalized in these 2 terms. So bias, sort of how well you are fitting the data essentially, and variance, like how your fits vary depending on, like if you sample different points from this distribution that you're trying to model. And it's true that sometimes if you naively build like a really large polynomial, for example, you can have low bias and high variance. Whereas if you build a small polynomial, maybe you have low variance and high bias. However, approaches like ensembling are actually a good way of getting low bias and low variance. And it turns out, actually building large neural nets are another way of getting both low bias and low variance. You actually have flexibility combined with a simplicity bias. And this is what's leading to good generalization, and it's sort of another perspective on double descent. I think here's here's the way I'll put the question is and I ran into this as a practitioner, you know, back in the day. So maybe I was just stuck in the hump of having, like, too many but not enough parameters, you know, to to get to the double descent phase. I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all. Right? Like, I mean, like, we kind of, you know, we'll talk about marginalization and kind of Bayesian Bayesian perspective on that. So, like, overfitting can be a real problem. So for example, you brought up, you know, conservation. It's like, well, if I'm doing a model and I don't enforce conservation of energy. And then as a result, I end up with some small parameters that cause, like, a little bit of feedback and increasing, you know, energy every single time a robot, you know, takes some action that can cause it to, like, spin out of control. Right? Where I actually did need it to conserve energy and not have that that positive feedback. So I guess I guess maybe we're still struggling with, you know overfitting can be a problem. It's a real problem. It's a known problem. How do we know if we're overfitting in a bad way? We maybe haven't don't have enough parameters. We're stuck in kind of the the area before we got to double descent. Or, like, how do you, in practice, avoid the actual consequences of bad overfitting? Right. So overfitting is real, absolutely. But the conventional wisdom about how we should approach it, I think, is fundamentally misguided. So and this is rooted in things like the bias variance trade off. Like, let's constrain our hypothesis space so that we can't have a bad fit to the data that will make bad predictions and so on. Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit. And interestingly, the prescription is almost the opposite of what people think it it perhaps should be in principle. Like, build a smaller model is usually the the prince the sort of the the prescription for avoiding overfitting.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence