High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Andrew Gordon Wilson: observation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“The reason is you're paying some sort of penalty, even if it's small, for deviating from that constraint.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
observation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. But how should we approach model construction? How can we embrace expressiveness without overfitting? And so when I say I want to embrace expressiveness, there is some some subtlety associated with that idea. And that basically means that it in some cases maybe we're wanting to represent lots of solutions, but we're assigning them almost 0 probability, but not 0 probability. So they're possible but not plausible solutions in our view. And then if the data is telling us something that like well actually we really should be paying attention to certain type of structure that maybe would surprise us, the model can actually respond to that. And if that structure isn't actually there, then your model isn't going to perform a lot worse than the model that is exactly constrained in those ways. In terms of how we should approach model construction in general, if you have a soft bias, a gentle encouragement towards certain types of constraints over others, quite often you can do as well as the perfectly constrained models. The reason is you're paying some sort of penalty, even if it's small, for deviating from that constraint. If So you can fit the data perfectly with the constraint, you'll just collapse down onto that model. And we noticed this in a work we had called residual pathway priors for soft equivariance constraints. So this was a Bayesian mechanism to have a distribution over solutions which would be concentrated in some way around certain types of equivariance constraints. Just for the sake of the audience, equivariance is a generalization of invariance. It basically means you have some transformation t, f of t x equals t of f of x, rather than f of t x equals f of x. Like a CNN, for example. So it it commutes in that case with the translation.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence