High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Andrew Gordon Wilson: belief

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
belief
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Join our Discord. I'd love to. Yeah. We actually have a yeah. I don't wanna get ahead of myself. We have an idea of course for how we can do this. So the the blog posts sort of compares this to end. So the entropy of the system is increasing, the Kolmogorov complexity is increasing. But the intuitive sophistication of that system is kind of non monotonic. Initially, it has low entropy, low sophistication, then sort of intermediate sophistication and entropy, and then again sort of like oh, sorry. And then high entropy and kind of low sophistication. And so you can think of this in machine in a machine learning context in terms of reasoning about the value of data. So if I sample data from like a uniform random uniform distribution, that's gonna be very incompressible. I'm gonna need to memorize it. It's sort of uncorrelated. This could be useless for learning a representation, for training my model. I could alternatively imagine some sort of sophisticated cellular automata problem, some sort of game of life problem with very sophisticated generalization rule or generation rules. That data actually could have an extraordinary amount of value for learning a representation. There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera. And so we've been thinking about kind of measures of information that might compartmentalize structural complexity and random complexity. And this will help us reason better about the value of data and developing priors sort of that are like Solomon of priors, might actually be more directly addressing the type of incompressibility that we're interested in. Yeah. I mean, so you said you said so many interesting things. I mean, of all, to Keith's point, you were saying that this isn't a penalty term yet your hypothesis is that neural networks implicitly do this kind of compression, which might be correlated or related to this Kormigrov complexity. So many folks just kind of they equate intelligence with compression. And when I spoke with David Krakauer, he took umbrage of that. He said, you know, compression is a component of of intelligence, but there are so many other things going on as well. And certainly, when we look at things like the ARC challenge, there are so many possible solutions. So a naive heuristic of just selecting the simplest program isn't always the best thing to do. There are many possible selections that you could make. And, you know, so so you you you demonstrated this upper bound which used this complex detail. And in a sense, that's saying that it could be no worse than this rather than it could be but it could actually be so much better. And our empirical experience of deep learning models is that they they seemed, you know, we call it shortcut learning basically. Know, they they they seem to have found some superficial generalizing thing which does all the things you said when you mentioned the no free lunch theorem. So you could take a CNN and you could use it on tabular data. You could take a transformer and you could use it on audio data. So it's almost like what we've seen is that we've we've hit this, for want of a better word, local minimum. And it feels like we need something more to get to the real understanding of some of these problems. Does does does that make sense?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence