Evidence receipt / belief
Published · transcript-backedAndrew Gordon Wilson: belief
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Join our Discord. I'd love to. Yeah. We actually have a yeah. I don't wanna get ahead of myself. We have an idea of course for how we can do this. So the the blog posts sort of compares this to end. So the entropy of the system is increasing, the Kolmogorov complexity is increasing. But the intuitive sophistication of that system is kind of non monotonic. Initially, it has low entropy, low sophistication, then sort of intermediate sophistication and entropy, and then again sort of like oh, sorry. And then high entropy and kind of low sophistication. And so you can think of this in machine in a machine learning context in terms of reasoning about the value of data. So if I sample data from like a uniform random uniform distribution, that's gonna be very incompressible. I'm gonna need to memorize it. It's sort of uncorrelated. This could be useless for learning a representation, for training my model. I could alternatively imagine some sort of sophisticated cellular automata problem, some sort of game of life problem with very sophisticated generalization rule or generation rules. That data actually could have an extraordinary amount of value for learning a representation. There was a paper that looked at something briefly like this called Intelligence at the Edge of Chaos, and I think there are other results like that that are coming out that you actually might want to train your models on this data with a lot of structural complexity, even if you really want in the end the model has some kind of outcomes, razor bias, etcetera. And so we've been thinking about kind of measures of information that might compartmentalize structural complexity and random complexity. And this will help us reason better about the value of data and developing priors sort of that are like Solomon of priors, might actually be more directly addressing the type of incompressibility that we're interested in. Yeah. I mean, so you said you said so many interesting things. I mean, of all, to Keith's point, you were saying that this isn't a penalty term yet your hypothesis is that neural networks implicitly do this kind of compression, which might be correlated or related to this Kormigrov complexity. So many folks just kind of they equate intelligence with compression. And when I spoke with David Krakauer, he took umbrage of that. He said, you know, compression is a component of of intelligence, but there are so many other things going on as well. And certainly, when we look at things like the ARC challenge, there are so many possible solutions. So a naive heuristic of just selecting the simplest program isn't always the best thing to do. There are many possible selections that you could make. And, you know, so so you you you demonstrated this upper bound which used this complex detail. And in a sense, that's saying that it could be no worse than this rather than it could be but it could actually be so much better. And our empirical experience of deep learning models is that they they seemed, you know, we call it shortcut learning basically. Know, they they they seem to have found some superficial generalizing thing which does all the things you said when you mentioned the no free lunch theorem. So you could take a CNN and you could use it on tabular data. You could take a transformer and you could use it on audio data. So it's almost like what we've seen is that we've we've hit this, for want of a better word, local minimum. And it feels like we need something more to get to the real understanding of some of these problems. Does does does that make sense?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.