Evidence receipt / prediction
Published · transcript-backedTim Scarfe: prediction
10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
“Just to kind of play that back just so that everyone The idea is that there is I mean, we're talking about grammar here. But more broadly, we think that there are structured generative processes in the world.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 10 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…r, allowed us to compute any correlation in this model. And from that, we could actually understand how they are learned by machines. And indeed if you have a poor machine like a very shallow network, 1st of all, even those models would be essentially unlearnable. Okay, so there would be a lot to say and then maybe I can come back to that in high dimension, it's very hard to learn. But if you have a deep architecture, what we find is precisely the opposite of what you're saying. The reason why deep architecture cancels those tasks is precisely because they understand just like the physicists understood about pressure, velocity, field. They understand from data, from the statistics of data this hidden hierarchy. Otherwise, they would never be able to do this job. So they understand that there is this hidden hierarchy. And from it, they can perform the task that you want. Just to kind of play that back just so that everyone The idea is that there is I mean, we're talking about grammar here. But more broadly, we think that there are structured generative processes in the world. So, can think of those as being some kind of a constrained generative model. So, we're talking about syntax here. And when we do machine learning, we look at the output of that generative model. And the learning process ideally should be not to memorize the raw output but it should be to understand abstractly the model which generated it. Because creativity is about respecting the deep structure and the constraints. If you have the structure you can go on and generate many, many more things and you obey the rules, you're creative and So all the rest of you're saying that you've done experiments. So you've created a mathematical generative model which is parameterized. You can have as much depth as you want. And you found on shallow networks that basically it wasn't really learning any of the abstract structure. But when you have deep networks, it was. Yeah, exactly. So maybe I can put it into the context of creativity and this discussion we just had with Shamsky. So there is a question creativity. I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraint syntactic rules that the child would never have heard before. And And Chomsky gave this poverty of stimulus argument arguing that it was actually impossible to learn to become creative from example. And essentially, okay, this would be a very crude way of summarizing his argument. But I described the fact that you have the sort of generative tree like rich context free grammars assuming that that really capturing the fact that the world has a hierarchy of abstract concepts. And but you have other possible generative grammar. Some are much simpler. 1 is called regular grammar. So this will be a caricature but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chomsky's argument is to say that, well, even if you give me 1,000,000 sentences, okay, I can fit those sentences by a context free grammar, but I can also feed them by a much simpler I mean, it's a regular grammar, simpler in this classification, but to feed those sentences, would have to be awfully complicated, many, many rules. And that nativism and not racism, big debates on this question. And again, we felt like we want to address this question as physicists. So in our idealized world where the true world is Yarky Kone and as is Yarky, can machine learn to be creative or not? And what we find is that if you have a shallow network, what Chomsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those coarse grained variables. This hierarchical architecture leads very easily to some iterative calculation. And so what we we found is that indeed you can learn to be creative by having being exposed to a very small number of sentences. So let's say in our model, if d is in the size of the sentence, the number of sentences is huge with d, exponential in d. But the number of sentences you need to see to be creative is only polynomial in d. So these models are really a counter example to his argument. And ultimately it comes from the fact that machines have strong implicit bias. They are not comparing equally in an equal fashion all hypothesis. And if you're deep, you learn hierarchies. So that's a counter example to that. So in some sense, I would argue that in terms of what needs to be innate, if you have deep architecture, that does lots of that. This is not to say that so I'm arguing against Chomsky's argument. It doesn't mean that…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.