High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Matthieu Wyart: evaluation

10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

“It doesn't mean that what he inferred is incorrect. It's not because I think an argument is incorrect that the statement is incorrect.”

— Matthieu Wyart

Source trail

Everything needed to verify it.

Speaker
Matthieu Wyart
Attribution
Verified speaker
Claim type
evaluation
Recorded
10 Aug 2026
Publisher
Machine Learning Street Talk

Transcript context

…it was. Yeah, exactly. So maybe I can put it into the context of creativity and this discussion we just had with Shamsky. So there is a question creativity. I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraint syntactic rules that the child would never have heard before. And And Chomsky gave this poverty of stimulus argument arguing that it was actually impossible to learn to become creative from example. And essentially, okay, this would be a very crude way of summarizing his argument. But I described the fact that you have the sort of generative tree like rich context free grammars assuming that that really capturing the fact that the world has a hierarchy of abstract concepts. And but you have other possible generative grammar. Some are much simpler. 1 is called regular grammar. So this will be a caricature but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chomsky's argument is to say that, well, even if you give me 1,000,000 sentences, okay, I can fit those sentences by a context free grammar, but I can also feed them by a much simpler I mean, it's a regular grammar, simpler in this classification, but to feed those sentences, would have to be awfully complicated, many, many rules. And that nativism and not racism, big debates on this question. And again, we felt like we want to address this question as physicists. So in our idealized world where the true world is Yarky Kone and as is Yarky, can machine learn to be creative or not? And what we find is that if you have a shallow network, what Chomsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those coarse grained variables. This hierarchical architecture leads very easily to some iterative calculation. And so what we we found is that indeed you can learn to be creative by having being exposed to a very small number of sentences. So let's say in our model, if d is in the size of the sentence, the number of sentences is huge with d, exponential in d. But the number of sentences you need to see to be creative is only polynomial in d. So these models are really a counter example to his argument. And ultimately it comes from the fact that machines have strong implicit bias. They are not comparing equally in an equal fashion all hypothesis. And if you're deep, you learn hierarchies. So that's a counter example to that. So in some sense, I would argue that in terms of what needs to be innate, if you have deep architecture, that does lots of that. This is not to say that so I'm arguing against Chomsky's argument. It doesn't mean that , I would argue that in terms of what needs to be innate, if you have deep architecture, that does lots of that. This is not to say that so I'm arguing against Chomsky's argument. It doesn't mean that what he inferred is incorrect. It's not because I think an argument is incorrect that the statement is incorrect. I don't want to imply that our brain is just a deep net and that they are not much smarter mechanisms to learn much better. And actually the brain can learn with 100,000 times less exposition towards than those machines. So I think there are lots of question about how the brain work and they are fascinating. I think it's possible that you're both correct. So Chomsky even back in the 50s and he wasn't the 1st to do this but he came up with these very basic transformative rules that could be composed together. And that went on for quite a while but they realised there were lots of problems and edge cases. And then eventually the minimalist programme came out and it was even more parsimonious. It was like move and merge. And that's very, very abstract because I agree with you. These networks clearly have syntactic competence, which means they absolutely are creative because they absolutely can create novel grammatical sentences. But in a broad setting though, they're not creative because they don't understand many other abstractions in the world that we do, which is why we need to prompt them to be creative. That they can render an image of a dog but they're not creative enough to know what an interesting and worldly coherent image of a dog is. But another thing I wanted to get to is this is touching on the idea that networks as we train them for longer and they get bigger and deeper, they start to factorize. So because we have this notion that they have these fractured and tangled representations. They understand things at a shallow level. They understand some things and not other things. I spoke with Tom McGrath at Goodfire. A big mechanistic interpretability company. And he's been studying networks for years. And he says as they get bigger and bigger they become more factorized. And he actually believes they're converging towards some kind of natural factorization. But at the moment, it's a bit weird that they have some factorization and then they have some fractionation in other areas. But you kind of see where I'm going with this because we know that a neural network is a finite state automata. And language in Chomsky sense, it's somewhere between context free and context sensitive. He wasn't super specific about that. But we know as a mathematical fact that it's not a generative grammar in the abstract way Chomsky was describing. But it is still a coherent generative grammar in some lower sense. Into models,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence