High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Matthieu Wyart: belief

10 Aug 2026 Machine Learning Street Talk AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

“What we argue is very important to look at and that we could finally measure with LLMs or other architecture and we find consistent result is what is the entropy left after a sentence of N token.”

— Matthieu Wyart

Source trail

Everything needed to verify it.

Speaker
Matthieu Wyart
Attribution
Verified speaker
Claim type
belief
Recorded
10 Aug 2026
Publisher
Machine Learning Street Talk

Transcript context

…And we could okay. So actually we have some confidence on those statements also because they make non trivial predictions. So for example, you would predict that as you train a diffusion more and more to generate text that initially it would be random. If you don't have data, it's regenerating crap. As you increase the number of data, it should start to form coherent words then later on with more data coherent group of words and then coherent full sentences. And this is a prediction of those models that coherence, the sort of coherence of the context should steadily increase as you increase the number of data that we could actually check for diffusion models. And also checking our theory of scaling laws of next token prediction. About a year ago, you had a paper about scaling laws as well. Tell me about that. As you increase the number of data or you increase how much compute you put or you increase the number of parameters, your performance steadily improves. And this observation by Kaplan and others had a huge impact for all of us because it drove the tech companies to just invest more and build maybe nuclear plants. So it has a huge technological impact, but it's a bit embarrassing for us here is that essentially it's not understood at all. You know, those getting those quantitatively, they have exponents in them. For example, describing how well you perform better if you multiply the number of data by 10. And there was very limited understanding on that question. And so yeah. So just a few months back with Francisco Cagneta, Alain Raventos and Suya Ganguly, we proposed a theory for this problem inspired by those synthetic world I told you about, but detaching sort of essence of the lesson we learned from those models to really make quantitative prediction for natural languages. And essentially, we the theory predicts that there is a simple recipe to force those exponents to extract those exponents. And and essentially this is saying, again, what's underlying it is the fact that if you give me more data, I can learn more abstract concept and that's longer range leads to longer range correlation. But at the end of the day, the 2 quantities you need to measure is 1, the fact that words or tokens are correlated and that this correlation, it was well known before us that this correlation decreases as a Paulo of the distance between those 2 words. From that you can measure exponents and they depend on the language you look at. As your data set, you can measure them. And then there is another key quantity we argue, which is related to the entropy of text. Entropy of text has been discussed already by Shannon in the 50s. It's a beautiful question. So essentially, the entropy is a log of the number of possible words that you would have at 1 location in average. What we argue is very important to look at and that we could finally measure with LLMs or other architecture and we find consistent result is what is the entropy u would have at 1 location in average. What we argue is very important to look at and that we could finally measure with LLMs or other architecture and we find consistent result is what is the entropy left after a sentence of N token. If you see N token, the more token you see the least possibility you have there. What is the entropy of that? And in the toy model of its Apollo and in real life, it's also Apollo that is found. And so essentially what we argue that with those 2 exponents, you can combine them in a way that we specify to get the training curve exponent of LLMs acting on those natural languages. And it works very well. So we got very excited with that. Also I have to say that in addition, it's making non trivial prediction in terms of how the loss should depend on the context that you give it, the context and also the number of data. So it's a function of 2 variables and we predict that it should bend in a very specific way and we could test it and we also observe it. And can you give me some interest on that? Mean, it's often when we speak about Shannon, a graph that people often see that during a sentence every single word reduces entropy. And we're now talking almost at the population scale that when we have a huge corpus of text, entropy is going down and down. What does that mean? Does that mean that the problem is getting easier over time? Does it mean that the models will just continue to get better or maybe there'll be some phase change? What's going happen? So I should be extremely careful 1st of all because…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence