← All source episodes Machine Learning Street Talk / episode intelligence
Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
19 Sept 2025 24 published claims 3 attributable people
Speakers in the public record
Claim mix
belief 10evaluation 8recommendation 3uncertainty 1observation 1prediction 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
24 published records
“So I think the bias variance trade off is an incredible misnomer. There doesn't actually have to be a trade off.”
- Publisher
- Machine Learning Street Talk
“Whereas in fact, as we build bigger models, we often actually start to alleviate overfitting. And double descent is just a great example of this because that first ascent is is from overfitting the data.”
- Publisher
- Machine Learning Street Talk
“I I sort of joke sometimes that if I hadn't met the airline passenger dataset, I don't know really what my life would be like now because it's it's driven so much of my research.”
- Publisher
- Machine Learning Street Talk
“There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera.”
- Publisher
- Machine Learning Street Talk
“Do we understand how the brain works? And this is sort of a conventional sort of like approach to science where the theory is really the quantity of primary interest and the applications of course are important, but they're not primarily why like no individual application is primarily why we care about the theory.”
- Publisher
- Machine Learning Street Talk
“I think although the field has made an extraordinary amount of empirical progress towards building more performant machine learning systems, We're still at early stages of understanding, you know, what principles should we broadly embrace when we're approaching our own problems.”
- Publisher
- Machine Learning Street Talk
“Absolutely. And I think being able to do certain things while we're finding might surprisingly relate to doing other things very well.”
- Publisher
- Machine Learning Street Talk
“Hopefully, they'll be able to go back and read not just my work, but like work that's been done in this space and think, okay, this is useful to me in thinking about how to approach some of these questions. And so, in this respect, I would say I'm a scientist and I try to combine classical theory with empiricism towards understanding model behavior.”
- Publisher
- Machine Learning Street Talk
“I mean, I think that's that's the type of reorganization is is what's happening in in double descent and probably grokking too.”
- Publisher
- Machine Learning Street Talk
“Like the models get both more expressive and they have a stronger simplicity bias. And so I think you can expand these 2 things together in some sense, and this is how you can avoid say overfitting and other sorts of issues with not achieving very good generalization.”
- Publisher
- Machine Learning Street Talk
“I think 1 of the most surprising findings in that paper was that convolutional neural nets, which were clearly designed for image recognition, so they have locality and translation equivariance and so on, provably have inductive biases for tabular data shaped as And an the only possible reason that could be the case is because they both sort of share this bias for low Kolmogorov complexity.”
- Publisher
- Machine Learning Street Talk
“I think that's a reasonable analogy. I would also add that we can't get away from making assumptions.”
- Publisher
- Machine Learning Street Talk
“Just a just a straight line. And I think was the second 1 something like 10 parameters and 10,000 parameters was the third 1.”
- Publisher
- Machine Learning Street Talk
“The reason is you're paying some sort of penalty, even if it's small, for deviating from that constraint.”
- Publisher
- Machine Learning Street Talk
“Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit.”
- Publisher
- Machine Learning Street Talk
“Like, once a certain number of people believe something, it's very, very hard to change their minds no matter what you say. And I think as a consequence, we've been in all sorts of local minima in machine learning and AI research because we haven't been able to get unstuck from these erroneous beliefs.”
- Publisher
- Machine Learning Street Talk
“A, because it's not an honest representative of our representation of our beliefs to to have those hard constraints, and b, because we see in practice that when we do have these expressive models with simplicity biases, they're much more adaptive.”
- Publisher
- Machine Learning Street Talk
“I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all.”
- Publisher
- Machine Learning Street Talk
“Because that means there are gonna be many more different settings of parameters that are consistent with what we observe, and we're just betting everything on 1 of them.”
- Publisher
- Machine Learning Street Talk
“Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now.”
- Publisher
- Machine Learning Street Talk
“Where like you have outliers but they're different each time. So just training on the outliers isn't gonna be useful because they're gonna be new outliers that look very different from those outliers.”
- Publisher
- Machine Learning Street Talk
“There are very subtle issues, think, with some of the ways that this can be done, and so we wrote a paper all about that. But largely speaking, it's, you know, something that I think people should acquaint themselves with because it it really is sort of getting at something very fundamental and it's has extraordinary practical value.”
- Publisher
- Machine Learning Street Talk
“ence on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. So this is really useful in practice, but it's very hard to kind of go beyond the useful approximations I think we developed without doing a significantly without making sort of a really significant, you know, 10 year kind of style moonshot landing kind of investment in those directions, which I think is worth making.”
- Publisher
- Machine Learning Street Talk
“Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems.”
- Publisher
- Machine Learning Street Talk