Evidence receipt / prediction
Published · transcript-backedJeremy Howard: prediction
3 Mar 2026 Machine Learning Street Talk "Vibe Coding is a Slot Machine" - Jeremy Howard
“You know, and because deep learning models are universal learning machines, you know, and we had a universal way to train them, I figured if if we get the data right and if the hardware is good enough, then in theory, we ought to be able to build that next word predicting machine, which ought to implicitly build a hierarchical structural understanding of the things that are being described by the text that it is learning to predict?”
Source trail
Everything needed to verify it.
- Speaker
- Jeremy Howard
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 3 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. I mean, I guess there's 2 things there. Mean, of all, we were kind of hinting at this notion of almost Goodhart's Law, the shortcut rule that you get exactly what you optimize for at the cost of everything else. But that doesn't seem to be the case because we can optimize for perplexity in the case of language models. And as you say, what seems to happen is we're getting into the distributional hypothesis here a little bit. So you know the word by the company it keeps. So when we have an incredible amount of associative data, it might be masked auto prediction or any of these things like that, the model seems to build something that we might call an understanding. Well, I I have always thought of it as a hierarchy of of abstractions, You know, it it it needs if it's gonna predict, you know, if the document is here was the, you know, opening that, you know, that Bobby Fischer used, and has chess notation to predict the next thing, it needs to know something about chess notation, or at least openings. If it's like, you know, and this was vetoed by the 1956 US president, comma, you need to know it's bit like, you don't even you don't just need to know who the president was, but the idea that there are presidents. And therefore, the idea that there are leaders, and therefore, the idea that there are groups of people who have hierarchies, and therefore, that there are people, and therefore, that there are objects, and like, you can't predict the next word of a sentence well without knowing all of these things. So that knowing my hypothesis for why I created ULMFiT was to say it would end to to to compress that as well as possible to get that knowledge, it would have to create these abstractions, these hierarchies of abstractions somewhere deep inside its model. Otherwise, how could it possibly do a good job of predicting the next word? You know, and because deep learning models are universal learning machines, you know, and we had a universal way to train them, I figured if if we get the data right and if the hardware is good enough, then in theory, we ought to be able to build that next word predicting machine, which ought to implicitly build a hierarchical structural understanding of the things that are being described by the text that it is learning to predict? I think that they can know in quite a, you know, they they know in quite a superficial way. So there's a myriad of surface statistical relationships, and they generalize extraordinarily well.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.