High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Andrew Gordon Wilson: prediction

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
prediction
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…But what if what if I'm a researcher who like, I just wanna I want this simplicity bias to go to 11, but I can't afford more parameters. Is there something I can tweak in my objective function to just push me a little bit more towards simplicity without breaking things? Yes. So there are a lot of things you can do. I think and we'll talk about them. But the question of how can you more elegantly encode this compression bias beyond just making the model bigger is really a fascinating open research question. So my contention, and I might be wrong, is that in many cases, a 7,000,000,000 parameter model is not doing better than a 1,000,000,000 parameter model primarily because it's more expressive. It's actually because of the simplicity bias. And perhaps things like knowledge distillation can help make this argument. If it's possible to really distill a large model into a much smaller model, then it means that the small model has some setting of its parameters that can provide a good approximation to the large model. It's just not able to find those parameters when trained directly on the data. It needs the help from the teacher model. And so the question is, can we build the 1,000,000,000 parameter model with some sort of like explicit bias that would have otherwise come about through scale in the 7,000,000,000 parameter model and find those solutions itself? I don't think we're anywhere close to being able to do that. I think it's Yeah. It's sort of an open research program. However, there are little things we can do that will help. So like there are ways of intervening in the training procedures. So like the stochastic weight averaging we discussed, that will help find a flatter, more compressible solution. There are of course, all sorts of regularizers that can be useful in certain instances. Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now. But then they have this sort of like more explicit compression bias that is interpretable and is getting us what building huge models is inelegantly giving us right now. Oh, that's interesting.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence