Evidence receipt / prediction
Published · transcript-backedMatthieu Wyart: prediction
10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
“Are we doing the wrong thing? And I'm very interested in, you know, should we predict in token space at a very low level or more should we train machine to predict abstractions.”
Source trail
Everything needed to verify it.
- Speaker
- Matthieu Wyart
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 10 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…I am Matthew Wyatt. I'm a full professor at John Hopkins University in The US and at EPFL in Switzerland. You know, those machines that can build new images that we've never seen before or say new sentences that were never heard before. Our brain seems to, you know, to to learn languages with 100,000 times less words than machines. Why is it so? Are we doing the wrong thing? And I'm very interested in, you know, should we predict in token space at a very low level or more should we train machine to predict abstractions. And so what we've been doing over the years is trying to build a framework based on physics that's really try to answer those different questions in a unified manner. Chomsky gave this poverty of stimulus argument arguing that it was actually impossible to learn to become creative from example. But if you have a deep architecture, there is a huge implicit bias to build those coarse grain variables. And so if you think about LLMs or diffusion models, the way they breed concepts, they emerge from statistics alone, just abstraction. They emerge out there in the data, they emerge. And those concept emerge if you group together configuration that predict similar context around them. And this is very pertinent because you've got a paper out basically saying that we should predict in the latent space, not the token space. So again, in those models, what we found is that…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.