High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Matthieu Wyart: evaluation

10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

“The reason why deep architecture cancels those tasks is precisely because they understand just like the physicists understood about pressure, velocity, field.”

— Matthieu Wyart

Source trail

Everything needed to verify it.

Speaker
Matthieu Wyart
Attribution
Verified speaker
Claim type
evaluation
Recorded
10 Aug 2026
Publisher
Machine Learning Street Talk

Transcript context

…nt level of abstraction. I mean it's also been argued for images. It's called pattern theory. Are very used to that in physics. In physics, let's say you take a liquid, I told you about liquid before. You can describe it at the level of atoms. But if you have billions of billions of billions of atoms in your glass, it's not going to help you describe the glass so well. And so what we did as physicists is to build coarse grain variables like pressure field, velocity, density, things like that. For this system it's sort of simpler because there is a single, I mean there are 2 level of description if you want essentially, this is a very microscopic and microscopic. Real data have layered multi scale level of description. I mean if you think about an image, you can think of it at a pixel level, at very low level. And at very high level, you could think of it as the caption that's describing what's on the image. And you have very intermediary steps. I mean at low level you could start to make from pixel edges and little geometrical figures. And at some point you could make eyes and nose and ears and understand that it makes a head that So you have many different level of description. And so the question you are asking, so that's the 1 we wanted to understand. And the 1st thing was how do we model this? Because if you look at the most complicated context free grammar, so those are the sort of models where which are based on trees that Shamsky introduced. Essentially the idea behind those models is that if you want to describe sort of linear objects that language text is, you can describe it by some underlying tree and you describe you have hidden variables living on those tree and you describe the way whereby those hidden variables can give rise to strings of hidden variables. So essentially you're describing a generative way to make sentences in a recursive fashion. But again, if you want to sort of feed those models to English, it's lots of it's very complicated. And so what is a good model? And again, it depends on the question you're asking. It's not an absolute, there is no absolute answer to that. But for the sort of question we're asking like the 1 you're asking which is in which sense do we build those cosmic variables? Well, we build context free grammar, some models of data where you have those trees underlying it. But like you as a physicist, it's a synthetic world, we build it. And then once we build it, the game we're playing is that we have to believe that it's rich enough. So in our case we want it to capture the fact that there is some hierarchical hidden structure to the world. So this we capture, but we want to make it tractable. So maybe it's a bit technical. In this case, we started the 1st models essentially at a tree that was frozen in geometry and production rules which is how latent gives rise to a string were randomly chosen. Finally randomness, although it's counterintuitive in physics often make things simpler, allowed us to compute any correlation in this model. And from that, we could actually understand how they are learned by machines. And indeed if you have a poor machine like a very shallow network, r, allowed us to compute any correlation in this model. And from that, we could actually understand how they are learned by machines. And indeed if you have a poor machine like a very shallow network, 1st of all, even those models would be essentially unlearnable. Okay, so there would be a lot to say and then maybe I can come back to that in high dimension, it's very hard to learn. But if you have a deep architecture, what we find is precisely the opposite of what you're saying. The reason why deep architecture cancels those tasks is precisely because they understand just like the physicists understood about pressure, velocity, field. They understand from data, from the statistics of data this hidden hierarchy. Otherwise, they would never be able to do this job. So they understand that there is this hidden hierarchy. And from it, they can perform the task that you want. Just to kind of play that back just so that everyone The idea is that there is I mean, we're talking about grammar here. But more broadly, we think that there are structured generative processes in the world. So, can think of those as being some kind of a constrained generative model. So, we're talking about syntax here. And when we do machine learning, we look at the output of that generative model. And the learning process ideally should be not to memorize the raw output but it should be to understand abstractly the model which generated it. Because creativity is about respecting the deep structure and the constraints. If you have the structure you can go on and generate many, many more things and you obey the rules, you're creative and So all the rest of you're saying that you've done experiments. So you've created a mathematical generative model which is parameterized. You can have as much depth as you want. And you found on shallow networks that basically it wasn't really learning any of the abstract structure. But when you have deep networks,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence