High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Andrew Gordon Wilson: belief

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“Like the models get both more expressive and they have a stronger simplicity bias. And so I think you can expand these 2 things together in some sense, and this is how you can avoid say overfitting and other sorts of issues with not achieving very good generalization.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
belief
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…So that has to be true. So how do we know in reality when we've stepped too far? It's a great question. I think we have to be careful about what we mean when we say a complex model. So I think most people would not consider Gaussian processes with an RBF covariance function, just the standard covariance function kernel that's often used to be a complex model. But it's highly expressive, so it's more expressive than any neural network can fit in memory. It just has very very strong preferences for certain types of solutions over others. This enables it to be extraordinarily data efficient. So 1 of the main use cases these days for Gaussian processes is in something called Bayesian optimization, where you're trying to maximize some sort of black box objective. So it's not something you have a closed form expression for, like it could be generalization performance of a neural net as a function of some of its hyper parameters for instance, or some really costly physical simulation as a function of some parameters, and you basically want to query this objective as few times as possible in order to achieve a good result. Gaussian processes are an amazing surrogate model for this objective, and you use the uncertainty to do the exploration efficiently. So I think that you can have expressive models that aren't necessarily complex. They still have very strong simplicity biases, very strong preferences for certain types of solutions over others, but at the same time they're representing a wide array of possible solutions to the problem. And I think that's how you can kind of reconcile what Radford Neil was saying with what we see in practice around things like overfitting. So I think there's a common misconception that the expressiveness of a model and its inductive biases are at odds with each other. The more expressive the model, the weaker its assumptions in some sense. Like the the fewer inductive biases it has, the less data efficient it will be, etcetera. And what we found, has been quite exciting, is that the larger you make say big transformers, actually the stronger its inductive biases. Like the models get both more expressive and they have a stronger simplicity bias. And so I think you can expand these 2 things together in some sense, and this is how you can avoid say overfitting and other sorts of issues with not achieving very good generalization. And I think 1 of the clearest demonstrations of this is in a phenomenon called double descent. So double descent is this phenomenon where typically on the horizontal axis you have the expressiveness of the model, the number of units for example in each layer of a residual neural network. And on the vertical axis you have generalization error. And so initially generalization error decreases as you increase the expressiveness of the model and it's able to just fit the data better and capture more structure. And then it starts to decrease, it starts to go up, and that corresponds to some sort of overfitting. expressiveness of the model and it's able to just fit the data better and capture more structure. And then it starts to decrease, it starts to go up, and that corresponds to some sort of overfitting. And then it decreases again, and that's why it's called second double descent, because of that second descent. And in that second descent, all the models typically have about 0 training loss. So the training loss just keeps going down as you increase the expressiveness of the model until roughly the number of parameters equals the number of data points. And what that means is that the larger models in that second descent cannot be generalizing better because they're more flexible. They're all fitting the training data perfectly. It has to be that the larger models have some sort of bias which is enabling better generalization. It turns out that this is a simplicity bias, a compression bias that we can measure.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence