Evidence receipt / evaluation
Published · transcript-backedMatthieu Wyart: evaluation
10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
“We make prediction and we test them. So we predict how many data you need to learn how many different level of abstraction.”
Source trail
Everything needed to verify it.
- Speaker
- Matthieu Wyart
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 10 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah, it's such a tantalizing idea. And we'll get to the latent stuff just in a minute because that's also a great thing to talk about. But are you suggesting that there is some kind of natural factorization? So do you think that different networks perhaps with different architectures given the same data would almost converge towards the same logical factorization of the data, this hierarchical factorization. So in our dream world that we invented mathematically, it's true as long as the network are deep. So shallow networks that don't do anything but if you have a transformer or CNN and so on, they build essentially very similar abstraction. Actually those arguments also predict how many data you need to build them. So this is a sort of evidence. We make prediction and that's what we play as physicists. We make prediction and we test them. So we predict how many data you need to learn how many different level of abstraction. And we find similar in terms of scaling similar number of data independently of the architecture. So I still do think that different architecture are going to lead to slight difference and so on different circuit. But the big picture I think is yes, those abstractions actually really exist. And again, those set of configurations that predict a similar surrounding. And if they have a strong signal, if they have a strong predictive power, they are formed earlier. It's so tantalizing to think of them as being the 1 true abstractions. But we know as we said earlier, it's not like the merge operator that Chomsky was talking about. When we do mechanistic interpretability and look at how networks do addition for example that it's super weird that they're composing trigonometric functions together. They're not doing it the way we can. And maybe that's just a limitation of the architecture. Maybe if we had proper learnable Turing machines that they would converge higher up the abstraction tree. But suppose a tangential thing that we need to talk about is this curse of dimensionality. So there's always been this statistical law essentially that when we have high dimensions the number of data that you need to make it tractable increases exponentially. And there were all of these theories about why that's not the case. There's the manifold hypothesis. So the intrinsic dimension is lower. We spoke with Randall Bellistriero about this, he's got this spline theory of neural networks and he said that in high dimensions all data is extrapolation, there's no manifold, that's bullshit. It's actually doing this spline decomposition sensitive way. Lots of people have different ideas about this. Yes. But you're saying that this kind of emergent behavior is actually how it is tractable.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.