Evidence receipt / prediction
Published · transcript-backedMatthieu Wyart: prediction
10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
“I mean the only way you will extrapolate and have power to generalize is if you bring those points together, it means you have an exponentially larger number of data, you have more data than atoms in the universe.”
Source trail
Everything needed to verify it.
- Speaker
- Matthieu Wyart
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 10 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…And if they have a strong signal, if they have a strong predictive power, they are formed earlier. It's so tantalizing to think of them as being the 1 true abstractions. But we know as we said earlier, it's not like the merge operator that Chomsky was talking about. When we do mechanistic interpretability and look at how networks do addition for example that it's super weird that they're composing trigonometric functions together. They're not doing it the way we can. And maybe that's just a limitation of the architecture. Maybe if we had proper learnable Turing machines that they would converge higher up the abstraction tree. But suppose a tangential thing that we need to talk about is this curse of dimensionality. So there's always been this statistical law essentially that when we have high dimensions the number of data that you need to make it tractable increases exponentially. And there were all of these theories about why that's not the case. There's the manifold hypothesis. So the intrinsic dimension is lower. We spoke with Randall Bellistriero about this, he's got this spline theory of neural networks and he said that in high dimensions all data is extrapolation, there's no manifold, that's bullshit. It's actually doing this spline decomposition sensitive way. Lots of people have different ideas about this. Yes. But you're saying that this kind of emergent behavior is actually how it is tractable. Exactly. So actually that's before we started thinking about creativity, our 1st work was really trying to understand what sort of data structure allow DeepNet to actually perform. So it's exactly as you said, maybe I can say it again. So in physics we know that volume goes like a length to the exponent of the dimension. So in 3 d L cube, in 2 d L square if L is a length. So think about a large dimension. So if you think about an image, D maybe is a number of pixels at least naively. You If think about text, it may be the number of words in your sentence again naively. So those volumes are huge. They're exponential. They're exponential in large in the dimension. So what it means is that even if you give me 1,000,000,000,000 points because the volume is so huge, it's extremely far away from each other, extremely far away. And so if you have a machine that's just interpolating and now you ask a question about a new test point and you can prove mathematically that if the data has little structure like you're trying to learn to regress some functions that smooth, it's hopeless. I mean the only way you will extrapolate and have power to generalize is if you bring those points together, it means you have an exponentially larger number of data, you have more data than atoms in the universe. So it's just impossible. So to me this is completely fundamental question. And it's true that sometimes in the literature it starts to side by saying, okay, it's super naive to say that the dimension is a number of peaks that are in an image. In fact, it should really be that the data lie in a lower dimension manifold. And if you try to measure it, it's true that it lies in a lower dimension manifold. But this dimension is still large in dimension. And to me the big problem is that if this was the answer to this question, it will mean that very simple algorithm like kernel methods are ancestor of deep net or even shallow network. I mean, they do it perfectly. And if you give them a low dimensional manifold, you don't need to have any interesting architecture. But if you use those on text, I can tell you it fails lament. I mean, it's complete, does nothing. So the question is why do you need deep architectures? I mean and some of the thing you said is it does not answer that question. So that's really the question we are looking after. And so essentially the answer to that is that if the world is hierarchical, if it has hidden cost grain variables, those machines are super good to discover them. And they can discover them generically with a number of data that's not huge, polynomial in the dimension once again. And once they discover them, it's a sort of summary of what the data is. Instead of describing pixel by pixel, it's always a nose, ears and so on. So you're reducing the dimension of the problem essentially. And you can solve the curse of the matrix. I think this explanation has the advantage that I mean, xel, it's always a nose, ears and so on. So you're reducing the dimension of the problem essentially. And you can solve the curse of the matrix. I think this explanation has the advantage that I mean, whatever you explanation you come up with, it has to explain it needs to explain why you need deep networks.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.