Evidence receipt / belief
Published · transcript-backedTim Scarfe: belief
3 Mar 2026 Machine Learning Street Talk "Vibe Coding is a Slot Machine" - Jeremy Howard
“I I think there's a dichotomy though between continual learning, which is when we want to keep training the thing but maintain generality, versus fine tuning a thing to do something specific.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…there there was no received wisdom because nobody talked about it. No 1 cared, you know. It's just this sort of like, nearly no 1 cared. Transfer learning was just not something anybody thought about. And Rachel and I felt like it matters more than anything, you know, because only 1 person has to train a really big model once, and then the rest of us can all fine tune it. So we thought we just should learn how to do that really well. So we spend a lot of time just trying lots of things. But in the end, the intuition was pretty straightforward, and what intuitively seemed like it ought to work, basically always did work. Which is another big difference between how people still today tend to do ML research, is I think it's all about ablations, and you can't make any assumptions or guesses, and it's not at all true. I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave. I I think there's a dichotomy though between continual learning, which is when we want to keep training the thing but maintain generality, versus fine tuning a thing to do something specific. So there's always been this idea that, yes, you can make a model specific. You can bend it to your will. But you lose generality, and you kind of degrade the representation. So tell me about that. Yeah, there's some truth in that, although not as much as you might think. On the whole, the big problem is that people don't actually look at their activations and don't actually look at their gradients. So something we do in our software, in our fast AI software, is we have built into it this ability to to see in a glance what your entire network looks like. And once you've done it a few times, it just takes a couple of hours to learn, you can immediately see, oh, I I see. This is over trained or under trained or this layer that something went wrong. It's not a mystery, you know. So basically what happens is, for example, you end up with with dead neurons that go to a point where they they've got 0 gradient regardless of what you do with them. That often happens if they, you know, head off towards infinity. You can always fix that. So yeah, it's it's not as bad as people think by any means. Something that trains well for continuous learning, when done properly, can also be done well to train well for a particular task, if you're careful.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.