High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Machine Learning Street Talk / episode intelligence

Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

19 Sept 2025 24 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 10evaluation 8recommendation 3uncertainty 1observation 1prediction 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

24 published records

02 / evaluation

Whereas in fact, as we build bigger models, we often actually start to alleviate overfitting. And double descent is just a great example of this because that first ascent is is from overfitting the data.

“Whereas in fact, as we build bigger models, we often actually start to alleviate overfitting. And double descent is just a great example of this because that first ascent is is from overfitting the data.”
Publisher
Machine Learning Street Talk

03 / uncertainty

I I sort of joke sometimes that if I hadn't met the airline passenger dataset, I don't know really what my life would be like now because it's it's driven so much of my research.

“I I sort of joke sometimes that if I hadn't met the airline passenger dataset, I don't know really what my life would be like now because it's it's driven so much of my research.”
Publisher
Machine Learning Street Talk

04 / belief

There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera.

“There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera.”
Publisher
Machine Learning Street Talk

05 / belief

Do we understand how the brain works? And this is sort of a conventional sort of like approach to science where the theory is really the quantity of primary interest and the applications of course are important, but they're not primarily why like no individual application is primarily why we care about the theory.

“Do we understand how the brain works? And this is sort of a conventional sort of like approach to science where the theory is really the quantity of primary interest and the applications of course are important, but they're not primarily why like no individual application is primarily why we care about the theory.”
Publisher
Machine Learning Street Talk

06 / belief

I think although the field has made an extraordinary amount of empirical progress towards building more performant machine learning systems, We're still at early stages of understanding, you know, what principles should we broadly embrace when we're approaching our own problems.

“I think although the field has made an extraordinary amount of empirical progress towards building more performant machine learning systems, We're still at early stages of understanding, you know, what principles should we broadly embrace when we're approaching our own problems.”
Publisher
Machine Learning Street Talk

08 / belief

Hopefully, they'll be able to go back and read not just my work, but like work that's been done in this space and think, okay, this is useful to me in thinking about how to approach some of these questions. And so, in this respect, I would say I'm a scientist and I try to combine classical theory with empiricism towards understanding model behavior.

“Hopefully, they'll be able to go back and read not just my work, but like work that's been done in this space and think, okay, this is useful to me in thinking about how to approach some of these questions. And so, in this respect, I would say I'm a scientist and I try to combine classical theory with empiricism towards understanding model behavior.”
Publisher
Machine Learning Street Talk

10 / belief

Like the models get both more expressive and they have a stronger simplicity bias. And so I think you can expand these 2 things together in some sense, and this is how you can avoid say overfitting and other sorts of issues with not achieving very good generalization.

“Like the models get both more expressive and they have a stronger simplicity bias. And so I think you can expand these 2 things together in some sense, and this is how you can avoid say overfitting and other sorts of issues with not achieving very good generalization.”
Publisher
Machine Learning Street Talk

11 / belief

I think 1 of the most surprising findings in that paper was that convolutional neural nets, which were clearly designed for image recognition, so they have locality and translation equivariance and so on, provably have inductive biases for tabular data shaped as And an the only possible reason that could be the case is because they both sort of share this bias for low Kolmogorov complexity.

“I think 1 of the most surprising findings in that paper was that convolutional neural nets, which were clearly designed for image recognition, so they have locality and translation equivariance and so on, provably have inductive biases for tabular data shaped as And an the only possible reason that could be the case is because they both sort of share this bias for low Kolmogorov complexity.”
Publisher
Machine Learning Street Talk

15 / recommendation

Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit.

“Whereas, instead, I think if we just embrace the honest belief that there are many possible solutions even if they're not probable for any given problem, combined with this sort of simplicity bias, we won't tend to overfit.”
Publisher
Machine Learning Street Talk

16 / evaluation

Like, once a certain number of people believe something, it's very, very hard to change their minds no matter what you say. And I think as a consequence, we've been in all sorts of local minima in machine learning and AI research because we haven't been able to get unstuck from these erroneous beliefs.

“Like, once a certain number of people believe something, it's very, very hard to change their minds no matter what you say. And I think as a consequence, we've been in all sorts of local minima in machine learning and AI research because we haven't been able to get unstuck from these erroneous beliefs.”
Publisher
Machine Learning Street Talk

17 / evaluation

A, because it's not an honest representative of our representation of our beliefs to to have those hard constraints, and b, because we see in practice that when we do have these expressive models with simplicity biases, they're much more adaptive.

“A, because it's not an honest representative of our representation of our beliefs to to have those hard constraints, and b, because we see in practice that when we do have these expressive models with simplicity biases, they're much more adaptive.”
Publisher
Machine Learning Street Talk

18 / evaluation

I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all.

“I'm not sure. But, I mean, you know, the what I experienced is that, you know, having parameters in a model, even if they're very, very small because some you know, I put in some term in the objective function that forced them to be small is not the same thing as actually the simpler model that just didn't have them at all.”
Speaker
Keith Duggar
Publisher
Machine Learning Street Talk

19 / evaluation

Because that means there are gonna be many more different settings of parameters that are consistent with what we observe, and we're just betting everything on 1 of them.

“Because that means there are gonna be many more different settings of parameters that are consistent with what we observe, and we're just betting everything on 1 of them.”
Publisher
Machine Learning Street Talk

20 / prediction

Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now.

“Bayesian marginalization can be very helpful in encoding an automatic simplicity bias in what we do. And so there are all sorts of interventions that will help us with this, but I think the dream is that maybe we can embrace flexibility in 15, 20 years from now by building these non parametric models that like really do have an infinite number of parameters and are more expressive than any model we're using right now.”
Publisher
Machine Learning Street Talk

21 / recommendation

Where like you have outliers but they're different each time. So just training on the outliers isn't gonna be useful because they're gonna be new outliers that look very different from those outliers.

“Where like you have outliers but they're different each time. So just training on the outliers isn't gonna be useful because they're gonna be new outliers that look very different from those outliers.”
Publisher
Machine Learning Street Talk

22 / recommendation

There are very subtle issues, think, with some of the ways that this can be done, and so we wrote a paper all about that. But largely speaking, it's, you know, something that I think people should acquaint themselves with because it it really is sort of getting at something very fundamental and it's has extraordinary practical value.

“There are very subtle issues, think, with some of the ways that this can be done, and so we wrote a paper all about that. But largely speaking, it's, you know, something that I think people should acquaint themselves with because it it really is sort of getting at something very fundamental and it's has extraordinary practical value.”
Publisher
Machine Learning Street Talk

23 / evaluation

ence on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. So this is really useful in practice, but it's very hard to kind of go beyond the useful approximations I think we developed without doing a significantly without making sort of a really significant, you know, 10 year kind of style moonshot landing kind of investment in those directions, which I think is worth making.

“ence on materials engineering and things like that, I'll see lots and lots of talks using Bayesian optimization, Gaussian processes, neural networks with epidemic uncertainty representation, etcetera. So this is really useful in practice, but it's very hard to kind of go beyond the useful approximations I think we developed without doing a significantly without making sort of a really significant, you know, 10 year kind of style moonshot landing kind of investment in those directions, which I think is worth making.”
Publisher
Machine Learning Street Talk

24 / evaluation

Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems.

“Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems.”
Publisher
Machine Learning Street Talk
Search evidence