High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Matthieu Wyart: evaluation

10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

“My take is that those are more academic problems that you never encountered in in practice because sentences that loop for 50 times are extremely rare.”

— Matthieu Wyart

Source trail

Everything needed to verify it.

Speaker
Matthieu Wyart
Attribution
Verified speaker
Claim type
evaluation
Recorded
10 Aug 2026
Publisher
Machine Learning Street Talk

Transcript context

…I think it's possible that you're both correct. So Chomsky even back in the 50s and he wasn't the 1st to do this but he came up with these very basic transformative rules that could be composed together. And that went on for quite a while but they realised there were lots of problems and edge cases. And then eventually the minimalist programme came out and it was even more parsimonious. It was like move and merge. And that's very, very abstract because I agree with you. These networks clearly have syntactic competence, which means they absolutely are creative because they absolutely can create novel grammatical sentences. But in a broad setting though, they're not creative because they don't understand many other abstractions in the world that we do, which is why we need to prompt them to be creative. That they can render an image of a dog but they're not creative enough to know what an interesting and worldly coherent image of a dog is. But another thing I wanted to get to is this is touching on the idea that networks as we train them for longer and they get bigger and deeper, they start to factorize. So because we have this notion that they have these fractured and tangled representations. They understand things at a shallow level. They understand some things and not other things. I spoke with Tom McGrath at Goodfire. A big mechanistic interpretability company. And he's been studying networks for years. And he says as they get bigger and bigger they become more factorized. And he actually believes they're converging towards some kind of natural factorization. But at the moment, it's a bit weird that they have some factorization and then they have some fractionation in other areas. But you kind of see where I'm going with this because we know that a neural network is a finite state automata. And language in Chomsky sense, it's somewhere between context free and context sensitive. He wasn't super specific about that. But we know as a mathematical fact that it's not a generative grammar in the abstract way Chomsky was describing. But it is still a coherent generative grammar in some lower sense. Into models, 1st of all what we find is that as you train the machine more and more, those factorization or abstraction are created progressively. And if you have an immense machine and you give me more and more data, then you start to play with more and more abstract concepts. Those are the hardest to learn in this viewpoint. So that's my take on the being of your question. The 2nd question was, I think you're referring to the fact that transformers have a finite depth. It's a finite depth. And so I mean, if you are thinking about sentences where you have 50 compliments of sentences and that are looping like that, it may be very hard to reproduce and so on. But I think there's some more like academic. My take is that those are more academic problems that you never encountered in in practice because sentences that loop for 50 times are extremely rare. So so I'm not sure it's really I I I know that some people put a lot of attention on that. But I'm more as someone who likes to do empirical studies and so on. I don't know if those worries are actually relevant in practice. Would you agree though that there is a spectrum of abstraction? Yes, it's a different type of syntactic competence, but maybe it doesn't matter. There's the performance competence type thing. But an interesting example is I feed my entire code base into Claude code. And isn't it interesting that it doesn't really understand what my intentions were? So if I put it in a loop, I put it in an agent and I say just fix the bugs and just keep evolving this software. It doesn't respect the deep constraints now are my mental constraints. Like what was I trying to achieve with this? Right? What would I have done? So it has the syntactic competence. It knows how to write the code. And are you saying this is just a matter of the networks aren't good enough yet? When they do understand, when they have a theory of mind and they understand how the world works even more abstractly. Eventually we could just autonomously…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence