Evidence receipt / prediction
Published · transcript-backedJeremy Howard: prediction
3 Mar 2026 Machine Learning Street Talk "Vibe Coding is a Slot Machine" - Jeremy Howard
“I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave.”
Source trail
Everything needed to verify it.
- Speaker
- Jeremy Howard
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 3 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. Because the discriminative learning rate thing is interesting because I I think the received wisdom at the time was when you fine tune a model, if the learning rate is too high, you kind of blow out the representations. So I guess the wisdom was if if you don't have a really low learning rate, you'll just destroy the representations. I mean, there there was no received wisdom because nobody talked about it. No 1 cared, you know. It's just this sort of like, nearly no 1 cared. Transfer learning was just not something anybody thought about. And Rachel and I felt like it matters more than anything, you know, because only 1 person has to train a really big model once, and then the rest of us can all fine tune it. So we thought we just should learn how to do that really well. So we spend a lot of time just trying lots of things. But in the end, the intuition was pretty straightforward, and what intuitively seemed like it ought to work, basically always did work. Which is another big difference between how people still today tend to do ML research, is I think it's all about ablations, and you can't make any assumptions or guesses, and it's not at all true. I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave. I I think there's a dichotomy though between continual learning, which is when we want to keep training the thing but maintain generality, versus fine tuning a thing to do something specific. So there's always been this idea that, yes, you can make a model specific. You can bend it to your will. But you lose generality, and you kind of degrade the representation. So tell me about that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.