Evidence receipt / belief
Published · transcript-backedTim Scarfe: belief
3 Mar 2026 Machine Learning Street Talk "Vibe Coding is a Slot Machine" - Jeremy Howard
“Yeah. Because the discriminative learning rate thing is interesting because I I think the received wisdom at the time was when you fine tune a model, if the learning rate is too high, you kind of blow out the representations.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. So the how we do the fine tuning was something we had developed at Fast AI. So this is kind of year 1 of Fast AI. So this is still in our very early days. And 1 of the extremely controversial things we did was we felt that we should focus on fine tuning existing models, because we thought fine tuning was important. Some other folks were doing work contemporaneously with that, so Jason Yosinski did some really great research. I think it's during his PhD on how to fine tune models, and how good they can be, and some other folks in the in the computer vision world. We were, you know, amongst the first. There's a bunch of us kind of really investing in fine tuning. And so, yeah, we we felt that using a single learning rate to fine tune the whole thing all at once made no sense, because the different layers have different behaviors. This is 1 of the things Jason Yosinski's research also showed. We developed this idea of like, well, it's also way faster if you just train the last layer. Right? Because it only has to back prop the last layer. And then once that's pretty good, back prop the next the last 2 and then the last 3. And then we use something called discriminative learning rates. So different layers, we would give different learning rates to. And then another critical insight that no 1 realized for years, even though we told everybody, was that you actually have to fine tune every batch norm. So all the normalization layers, you do actually have to fine tune, because that's moving the whole thing up and down, know, changing its scale. So yeah, when you do that, you can often just fine tune the last layer or 2. And we found that actually with ULMFiT, although we did end up unfreezing all the layers, only the last 2 were really needed to get the close to state of the art result. So it like took like seconds. Yeah. Because the discriminative learning rate thing is interesting because I I think the received wisdom at the time was when you fine tune a model, if the learning rate is too high, you kind of blow out the representations. So I guess the wisdom was if if you don't have a really low learning rate, you'll just destroy the representations. I mean, there there was no received wisdom because nobody talked about it. No 1 cared, you know. It's just this sort of like, nearly no 1 cared. Transfer learning was just not something anybody thought about. And Rachel and I felt like it matters more than anything, you know, because only 1 person has to train a really big model once, and then the rest of us can all fine tune it. So we thought we just should learn how to do that really well. So we spend a lot of time just trying lots of things. But in the end, the intuition was pretty straightforward, and what intuitively seemed like it ought to work, basically always did work. Which is another big difference between how people still today tend to do ML research, is I think it's all about ablations, and you can't make any assumptions or guesses, and it's not at all true. I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.