Evidence receipt / recommendation
Published · transcript-backedJeremy Howard: recommendation
3 Mar 2026 Machine Learning Street Talk "Vibe Coding is a Slot Machine" - Jeremy Howard
“I'm a huge fan of taking a model that's incredibly flexible, and then making it more constrained, not by decreasing the size of architecture, but by adding regularization.”
Source trail
Everything needed to verify it.
- Speaker
- Jeremy Howard
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 3 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…just just sketch that out. I'm a huge fan of regularization. I'm a huge fan of taking a model that's incredibly flexible, and then making it more constrained, not by decreasing the size of architecture, but by adding regularization. So even that at the time was extremely controversial, but that was by no means a unique insight of of ours. So what Stephen Merity had done is he'd taken the extreme flexibility of an LSTM, a kind of the classic state for recurrent neural net towards which things are kind of gradually heading back towards nowadays, added added 5 different types of regularization. He added every type of regularization you can imagine, and then that was my starting point, was to say, now have a massively flexible deep learning model that can be as powerful as I want it to be, and it can also be as constrained as I need it to be, and then I needed a really big corpus of text. Funnily enough, this is also Steven. He had been at Common Crawl, and I think he helped or made the Wikipedia dataset. Then I realized actually the Wikipedia dataset made lots of assumptions. It had all these like, unk for unknown words, because it all assumed classic NLP approaches. So I redid the whole thing, created a new Wikipedia dataset. Now it's my general corpus. And then I used an AWD-LSTM, trained it so it's actually overnight. So for 8 hours on a gaming GPU, you know. Because I was at the University of San Francisco, we didn't have heaps of resources. Probably like a 20 80 TI or something, I suspect. And then the next morning when I woke up, I then it's the same 3 stage architecture that we do today. You know, pre training, mid training, post training. Mhmm. So then I figured, okay, now that I've trained something to predict the next word of Wikipedia, it must know a lot about the world. I then figured if I then fine tune it on a corpus specific, so what we could now call supervised fine tuning dataset, which in this case was the dataset of movie reviews. It would become especially good at predicting the next word of those, so I'd learn a lot about movies. Did that for like an hour, and then like a few minutes of fine tuning the downstream classifier, which was a classic academic data set. It's kind of considered the hardest 1, which was to take like 5,000 word movie reviews and to say like, is this a positive or negative sentiment? Which today is considered easy. But at that time, you know, the only things that did it quite well were highly specialized models that people wrote their whole PhDs on, and I beat all of their results, you know, 5 minutes later when it fine tuning that model. It was amazing. And the other interesting thing is this kind of methodology around how you do the fine tuning.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.