Evidence receipt / prediction
Published · transcript-backedAndrew Gordon Wilson: prediction
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…But what if what if I'm a researcher who like, I just wanna I want this simplicity bias to go to 11, but I can't afford more parameters. Is there something I can tweak in my objective function to just push me a little bit more towards simplicity without breaking things? Yes. So there are a lot of things you can do. I think and we'll talk about them. But the question of how can you more elegantly encode this compression bias beyond just making the model bigger is really a fascinating open research question. So my contention, and I might be wrong, is that in many cases, a 7,000,000,000 parameter model is not doing better than a 1,000,000,000 parameter model primarily because it's more expressive. It's actually because of the simplicity bias. And perhaps things like knowledge distillation can help make this argument. If it's possible to really distill a large model into a much smaller model, then it means that the small model has some setting of its parameters that can provide a good approximation to the large model. It's just not able to find those parameters when trained directly on the data. It needs the help from the teacher model. And so the question is, can we build the 1,000,000,000 parameter model with some sort of like explicit bias that would have otherwise come about through scale in the 7,000,000,000 parameter model and find those solutions itself? I don't think we're anywhere close to being able to do that. I think it's Yeah. It's sort of an open research program. However, there are little things we can do that will help. So like there are ways of intervening in the training procedures. So like the stochastic weight averaging we discussed, that will help find a flatter, more compressible solution. There are of course, all sorts of regularizers that can be useful in certain instances. Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now. But then they have this sort of like more explicit compression bias that is interpretable and is getting us what building huge models is inelegantly giving us right now. Oh, that's interesting.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.