Evidence receipt / commitment
Published · transcript-backedMatthieu Wyart: commitment
10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
“I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraint syntactic rules that the child would never have heard before.”
Source trail
Everything needed to verify it.
- Speaker
- Matthieu Wyart
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 10 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…you want. Just to kind of play that back just so that everyone The idea is that there is I mean, we're talking about grammar here. But more broadly, we think that there are structured generative processes in the world. So, can think of those as being some kind of a constrained generative model. So, we're talking about syntax here. And when we do machine learning, we look at the output of that generative model. And the learning process ideally should be not to memorize the raw output but it should be to understand abstractly the model which generated it. Because creativity is about respecting the deep structure and the constraints. If you have the structure you can go on and generate many, many more things and you obey the rules, you're creative and So all the rest of you're saying that you've done experiments. So you've created a mathematical generative model which is parameterized. You can have as much depth as you want. And you found on shallow networks that basically it wasn't really learning any of the abstract structure. But when you have deep networks, it was. Yeah, exactly. So maybe I can put it into the context of creativity and this discussion we just had with Shamsky. So there is a question creativity. I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraint syntactic rules that the child would never have heard before. And And Chomsky gave this poverty of stimulus argument arguing that it was actually impossible to learn to become creative from example. And essentially, okay, this would be a very crude way of summarizing his argument. But I described the fact that you have the sort of generative tree like rich context free grammars assuming that that really capturing the fact that the world has a hierarchy of abstract concepts. And but you have other possible generative grammar. Some are much simpler. 1 is called regular grammar. So this will be a caricature but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chomsky's argument is to say that, well, even if you give me 1,000,000 sentences, okay, I can fit those sentences by a context free grammar, but I can also feed them by a much simpler I mean, it's a regular grammar, simpler in this classification, but to feed those sentences, would have to be awfully complicated, many, many rules. And that nativism and not racism, big debates on this question. And again, we felt like we want to address this question as physicists. So in our idealized world where the true world is Yarky Kone and as is Yarky, can machine learn to be creative or not? And what we find is that if you have a shallow network, what Chomsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those coarse grained variables. This hierarchical architecture leads very easily to some iterative calculation. And so what we we found is that indeed you can learn to be creative by having being exposed to a very small number of sentences. So let's say in our model, if d is in the size of the sentence, the number of sentences is huge with d, exponential in d. But the number of sentences you need to see to be creative is only polynomial in d. So these models are really a counter example to his argument. And ultimately it comes from the fact that machines have strong implicit bias. They are not comparing equally in an equal fashion all hypothesis. And if you're deep, you learn hierarchies. So that's a counter example to that. So in some sense, I would argue that in terms of what needs to be innate, if you have deep architecture, that does lots of that. This is not to say that so I'm arguing against Chomsky's argument. It doesn't mean that , I would argue that in terms of what needs to be innate, if you have deep architecture, that does lots of that. This is not to say that so I'm arguing against Chomsky's argument. It doesn't mean that what he inferred is incorrect. It's not because I think an argument is incorrect that the statement is incorrect. I don't want to imply that our brain is just a deep net and that they are not much smarter mechanisms to learn much better. And actually the brain can learn with 100,000 times less exposition towards than those machines. So I think there are lots of question about how the brain work and they are fascinating.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.