High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Jeremy Howard

Published podcast speaker

Claims
39
Episodes
3
Shows
2
Named items
4

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

app / uses

Anki

“And I really noticed this that I used Anki, and because it was always scheduling my cards just before I was about to forget them, it was always incredibly hard work.”

Machine Learning Street Talk · 3 Mar 2026

Evidence receipt · Source ↗

tool / uses

Solveit

“It's like for me, when I use Solveit, it's the opposite of that experience you described with Claude Code. After a couple of hours, I feel energized and happy and fulfilled.”

Machine Learning Street Talk · 3 Mar 2026

Evidence receipt · Source ↗

app / uses

Solveit

“It's about, like, creating an environment where humans can grow and engage and share. It's like for me, when I use Solveit, it's the opposite of that experience you described with Claude Code.”

Machine Learning Street Talk · 3 Mar 2026

Evidence receipt · Source ↗

paper / uses

ULM Fit

“And then I read ULM Fit and turns out it did work. And so I did it, you know, bigger and it worked even better.”

Latent Space · 19 Oct 2023

Evidence receipt · Source ↗

Claim ledger

What Jeremy said.

39 transcript-backed records

01 / prediction

You know, and because deep learning models are universal learning machines, you know, and we had a universal way to train them, I figured if if we get the data right and if the hardware is good enough, then in theory, we ought to be able to build that next word predicting machine, which ought to implicitly build a hierarchical structural understanding of the things that are being described by the text that it is learning to predict?

“You know, and because deep learning models are universal learning machines, you know, and we had a universal way to train them, I figured if if we get the data right and if the hardware is good enough, then in theory, we ought to be able to build that next word predicting machine, which ought to implicitly build a hierarchical structural understanding of the things that are being described by the text that it is learning to predict?”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

03 / belief

I think in the end like IPykernel, I'm finding for example, it's just too big a piece, right? Because in the end, the the team that made the original IPykernel were not able to create a set of tests that correctly exercised it, and therefore real world downstream projects, including the original nb classic, you know, which is what IPykernel was extracted from, didn't work anymore.

“I think in the end like IPykernel, I'm finding for example, it's just too big a piece, right? Because in the end, the the team that made the original IPykernel were not able to create a set of tests that correctly exercised it, and therefore real world downstream projects, including the original nb classic, you know, which is what IPykernel was extracted from, didn't work anymore.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

04 / prediction

I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave.

“I find nearly everything that I expect to work almost always works first time, because I spend a lot of time building up those intuitions, that kind of understanding of how gradients behave.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

06 / belief

If there are risks with the current state of technology, I mean, I think some of them are the ones we've discussed, which is people enfeebling themselves by basically losing their ability to be to become more competent over time.

“If there are risks with the current state of technology, I mean, I think some of them are the ones we've discussed, which is people enfeebling themselves by basically losing their ability to be to become more competent over time.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

07 / commitment

Elon Musk said something a bit similar a few days ago, saying like, oh, LLMs will just spit out the machine code directly. We won't need libraries, programming languages.

“Elon Musk said something a bit similar a few days ago, saying like, oh, LLMs will just spit out the machine code directly. We won't need libraries, programming languages.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

08 / evaluation

The difference between pretending to be intelligent and actually being intelligent is entirely unimportant, as long as you're in the region in which the pretense is actually effective, you know. So so it's actually fine for a great many tasks that LLMs only pretend to be intelligent, because for all intents and purposes, it it it just doesn't matter until you get to the point where it can't pretend anymore.

“The difference between pretending to be intelligent and actually being intelligent is entirely unimportant, as long as you're in the region in which the pretense is actually effective, you know. So so it's actually fine for a great many tasks that LLMs only pretend to be intelligent, because for all intents and purposes, it it it just doesn't matter until you get to the point where it can't pretend anymore.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

09 / preference

I basically never have to use a debugger, because I basically never have bugs. And it's not because I'm a particularly good programmer, it's because I build things up small little steps, and each step works and I can see it working and I can interact with it.

“I basically never have to use a debugger, because I basically never have bugs. And it's not because I'm a particularly good programmer, it's because I build things up small little steps, and each step works and I can see it working and I can interact with it.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

10 / recommendation

I'm a huge fan of taking a model that's incredibly flexible, and then making it more constrained, not by decreasing the size of architecture, but by adding regularization.

“I'm a huge fan of taking a model that's incredibly flexible, and then making it more constrained, not by decreasing the size of architecture, but by adding regularization.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

11 / evaluation

They're really bad at software engineering. And then I think that's possibly always gonna be true, because, you know, we're we're asking them to often move outside of their training data, you know, if we're trying to build something that literally hasn't been built before and do it in a better way than has been done before, we're saying, like, don't just copy what was in the training data.

“They're really bad at software engineering. And then I think that's possibly always gonna be true, because, you know, we're we're asking them to often move outside of their training data, you know, if we're trying to build something that literally hasn't been built before and do it in a better way than has been done before, we're saying, like, don't just copy what was in the training data.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

12 / observation

I'm doing things that haven't been done before. And there's this weird thing, I don't know if you've ever seen it before, I see it but I see it multiple times every day, where the LM goes from being incredibly clever to like worse than stupid, like like not understanding the most basic fundamental premises about how the world works.

“I'm doing things that haven't been done before. And there's this weird thing, I don't know if you've ever seen it before, I see it but I see it multiple times every day, where the LM goes from being incredibly clever to like worse than stupid, like like not understanding the most basic fundamental premises about how the world works.”
Speaker
Jeremy Howard
Publisher
Machine Learning Street Talk

16 / observation

The point is, Sean, that once you've got your new model, if you distribute it as an adapter that sits on top of a quantized model that somebody's already downloaded, then it's a much smaller download for them.

“The point is, Sean, that once you've got your new model, if you distribute it as an adapter that sits on top of a quantized model that somebody's already downloaded, then it's a much smaller download for them.”
Speaker
Jeremy Howard
Publisher
Latent Space

17 / belief

We all understand kind of roughly why we're here, you know, we agree with the premises around, like, everything's too expensive, everything's too complicated, people are building too many vanity foundation models rather than taking better advantage of fine-tuning, like, there's this kind of general, like, sense of we're all on the same wavelength about, you know, all the ways in which current research is fucked up, and, you know, all the ways in which we're worried about centralization.

“We all understand kind of roughly why we're here, you know, we agree with the premises around, like, everything's too expensive, everything's too complicated, people are building too many vanity foundation models rather than taking better advantage of fine-tuning, like, there's this kind of general, like, sense of we're all on the same wavelength about, you know, all the ways in which current research is fucked up, and, you know, all the ways in which we're worried about centralization.”
Speaker
Jeremy Howard
Publisher
Latent Space

18 / recommendation

Once you've outgrown, or if you outgrow that, it's not like, okay, throw that all away and start again. And this like whole separate language that it's like this kind of smooth, gentle path that you can take step-by-step because it's all just standard web foundations all the way, you know.

“Once you've outgrown, or if you outgrow that, it's not like, okay, throw that all away and start again. And this like whole separate language that it's like this kind of smooth, gentle path that you can take step-by-step because it's all just standard web foundations all the way, you know.”
Speaker
Jeremy Howard
Publisher
Latent Space

19 / prediction

Sure. So, yeah, I mean, it was kind of uncomfortable because two days before Altman got fired, I did a small public video interview in which I said, I'm quite sure that OpenAI's current governance structure can't continue and that it was definitely going to fall apart.

“Sure. So, yeah, I mean, it was kind of uncomfortable because two days before Altman got fired, I did a small public video interview in which I said, I'm quite sure that OpenAI's current governance structure can't continue and that it was definitely going to fall apart.”
Speaker
Jeremy Howard
Publisher
Latent Space

20 / evaluation

Like I love it, I know a lot of people who didn't really know how to code, but they've created things because they use ChatGPT, but they don't really know how to maintain them or fix them or add things to them that ChatGPT can't do, because they don't really know how to code.

“Like I love it, I know a lot of people who didn't really know how to code, but they've created things because they use ChatGPT, but they don't really know how to maintain them or fix them or add things to them that ChatGPT can't do, because they don't really know how to code.”
Speaker
Jeremy Howard
Publisher
Latent Space

21 / preference

Yeah, I've, you know, that's been our kind of continuous message since we started Fast AI, is if you're training for random weights, you better have a really good reason, you know, because it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the only situation in which I think starting from random weights makes sense.

“Yeah, I've, you know, that's been our kind of continuous message since we started Fast AI, is if you're training for random weights, you better have a really good reason, you know, because it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the only situation in which I think starting from random weights makes sense.”
Speaker
Jeremy Howard
Publisher
Latent Space

22 / evaluation

I think one of the key things that's happened in all of these is everybody understands what Eric Gilliam, who wrote the second blog post in our series, the R&D historian, describes as a large yard with narrow fences.

“I think one of the key things that's happened in all of these is everybody understands what Eric Gilliam, who wrote the second blog post in our series, the R&D historian, describes as a large yard with narrow fences.”
Speaker
Jeremy Howard
Publisher
Latent Space

23 / commitment

You know, I know that if I didn't do it, then I would just get fired and the board would put in somebody else and the board knows if they don't do it, then their shareholders can sue them because they're not maximizing profitability or whatever.

“You know, I know that if I didn't do it, then I would just get fired and the board would put in somebody else and the board knows if they don't do it, then their shareholders can sue them because they're not maximizing profitability or whatever.”
Speaker
Jeremy Howard
Publisher
Latent Space

24 / evaluation

So we ended up, you know, trying to fix a whole lot of different things. And even as we did so, new regressions were appearing in like transformers and stuff that Benjamin then had to go away and figure out like, oh, how come flash attention doesn't work in this version of transformers anymore with this set of models and like, oh, it turns out they accidentally changed this thing, so it doesn't work.

“So we ended up, you know, trying to fix a whole lot of different things. And even as we did so, new regressions were appearing in like transformers and stuff that Benjamin then had to go away and figure out like, oh, how come flash attention doesn't work in this version of transformers anymore with this set of models and like, oh, it turns out they accidentally changed this thing, so it doesn't work.”
Speaker
Jeremy Howard
Publisher
Latent Space

25 / evaluation

I think people are starting to understand that treating the three ULM FIT steps of like pre-training, you know, and then the kind of like what people now call instruction tuning, and then, I don't know if we've got a general term for this, DPO, RLHFE step, you know, or the task training, they're not actually as separate as we originally suggested they were in our paper, and when you treat it more as a continuum, and that you make sure that you have, you know, more of kind of the original data set incorporated into the later stages, and that, you know, we've also seen with LLAMA3, this idea that those later stages can be done for a lot longer.

“I think people are starting to understand that treating the three ULM FIT steps of like pre-training, you know, and then the kind of like what people now call instruction tuning, and then, I don't know if we've got a general term for this, DPO, RLHFE step, you know, or the task training, they're not actually as separate as we originally suggested they were in our paper, and when you treat it more as a continuum, and that you make sure that you have, you know, more of kind of the original data set incorporated into the later stages, and that, you know, we've also seen with LLAMA3, this idea that those later stages can be done for a lot longer.”
Speaker
Jeremy Howard
Publisher
Latent Space

26 / evaluation

Actually, Karim's another great example of this, I mean, I already knew Karim very well because he was my best ever master's student, but it wasn't a surprise to me then when he then went off to create the world's state-of-the-art language model in Turkish on his own, in his spare time, with no budget, from scratch.

“Actually, Karim's another great example of this, I mean, I already knew Karim very well because he was my best ever master's student, but it wasn't a surprise to me then when he then went off to create the world's state-of-the-art language model in Turkish on his own, in his spare time, with no budget, from scratch.”
Speaker
Jeremy Howard
Publisher
Latent Space

27 / belief

You know, if we lock things down to the people that we think, you know, the elites that we think can be trusted to run it for us, yeah, I think all bets are off about where that leaves us as a society, you know.

“You know, if we lock things down to the people that we think, you know, the elites that we think can be trusted to run it for us, yeah, I think all bets are off about where that leaves us as a society, you know.”
Speaker
Jeremy Howard
Publisher
Latent Space

28 / belief

You know, it still requires kind of really understanding the GPU architecture and writing it in that kind of very CUDA-ish way. So yeah, I think, you know, if Mojo or something equivalent can really work well, we're going to see a lot more FlashAttentions popping up.

“You know, it still requires kind of really understanding the GPU architecture and writing it in that kind of very CUDA-ish way. So yeah, I think, you know, if Mojo or something equivalent can really work well, we're going to see a lot more FlashAttentions popping up.”
Speaker
Jeremy Howard
Publisher
Latent Space

29 / belief

You know, particularly in kind of an enterprise setting, I think there's a lot of like repetitive kind of processing that has to be done. It's a useful thing for coders to know about, because I think quite often you can like replace some thousands and thousands of lines of complex buggy code, maybe with a fine tune, you know.

“You know, particularly in kind of an enterprise setting, I think there's a lot of like repetitive kind of processing that has to be done. It's a useful thing for coders to know about, because I think quite often you can like replace some thousands and thousands of lines of complex buggy code, maybe with a fine tune, you know.”
Speaker
Jeremy Howard
Publisher
Latent Space

30 / recommendation

You know, work with vision, work with tables of data, work with kind of recommendation systems and collaborative filtering and work with text, because we felt like those four kind of modalities covered a lot of the stuff that, you know, are useful in real life.

“You know, work with vision, work with tables of data, work with kind of recommendation systems and collaborative filtering and work with text, because we felt like those four kind of modalities covered a lot of the stuff that, you know, are useful in real life.”
Speaker
Jeremy Howard
Publisher
Latent Space

32 / evaluation

You know, and I showed again through research that we demonstrated in our videos that you can do better than GANs, much faster and with much less data. And nobody cared because again, like if you want to get published, you write a GAN paper that slightly improves this part of GANs and this tiny field, you'll get published, you know.

“You know, and I showed again through research that we demonstrated in our videos that you can do better than GANs, much faster and with much less data. And nobody cared because again, like if you want to get published, you write a GAN paper that slightly improves this part of GANs and this tiny field, you'll get published, you know.”
Speaker
Jeremy Howard
Publisher
Latent Space

33 / preference

I try to make like, when I work with some domain, I try to make it like, I want to make it as enjoyable as possible for me to do that. So I always try to kind of like, like with GHAPI, for example, I think that GitHub API is incredibly powerful, but I didn't find it good to work with because I didn't particularly like the libraries that are out there.

“I try to make like, when I work with some domain, I try to make it like, I want to make it as enjoyable as possible for me to do that. So I always try to kind of like, like with GHAPI, for example, I think that GitHub API is incredibly powerful, but I didn't find it good to work with because I didn't particularly like the libraries that are out there.”
Speaker
Jeremy Howard
Publisher
Latent Space

34 / disagreement

Even though I originally created three-step approach that everybody now does, my view is it's actually wrong and we shouldn't use it. And that's because people are using it in a way different to why I created it.

“Even though I originally created three-step approach that everybody now does, my view is it's actually wrong and we shouldn't use it. And that's because people are using it in a way different to why I created it.”
Speaker
Jeremy Howard
Publisher
Latent Space

35 / prediction

Yeah. And then when I came across neural nets when I was about 20, you know, what I learned about the universal approximation theorem and stuff, and I started thinking like, oh, I wonder if like a neural net could ever get big enough and take in enough data to be a Chinese room experiment.

“Yeah. And then when I came across neural nets when I was about 20, you know, what I learned about the universal approximation theorem and stuff, and I started thinking like, oh, I wonder if like a neural net could ever get big enough and take in enough data to be a Chinese room experiment.”
Speaker
Jeremy Howard
Publisher
Latent Space

36 / evaluation

You know, it took a lot longer than it should have because I spent way longer in management consulting than I should have because I got caught up in that stupid rat race.

“You know, it took a lot longer than it should have because I spent way longer in management consulting than I should have because I got caught up in that stupid rat race.”
Speaker
Jeremy Howard
Publisher
Latent Space

37 / evaluation

5 has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know, it doesn't know who anybody is, it doesn't know about any movies, it doesn't really know anything about anything, like, because it's never read anything, you know, it was trained on a nearly entirely synthetic data set, which is designed for it to learn reasoning, and so it was a research project, and a really good one, and it definitely shows us a powerful direction in terms of what you can do with synthetic data, and wow, gosh, even these tiny models can get pretty good reasoning skills, pretty good math skills, pretty good coding skills, but I don't know if it's a model you could necessarily build on.

“5 has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know, it doesn't know who anybody is, it doesn't know about any movies, it doesn't really know anything about anything, like, because it's never read anything, you know, it was trained on a nearly entirely synthetic data set, which is designed for it to learn reasoning, and so it was a research project, and a really good one, and it definitely shows us a powerful direction in terms of what you can do with synthetic data, and wow, gosh, even these tiny models can get pretty good reasoning skills, pretty good math skills, pretty good coding skills, but I don't know if it's a model you could necessarily build on.”
Speaker
Jeremy Howard
Publisher
Latent Space

38 / preference

And then in step three, rather than fine tuning on a reasonably specific task classification, let's fine tune on a, on a RLHF task classification. And so that was really, that was really key, you know, so I was kind of like out of the NLP field for a few years there because yeah, it just felt like, I don't know, pushing uphill against this vast tide, which I was convinced was not the right direction, but who's going to listen to me, you know, cause I, as you said, I don't have a PhD, not at a university, or at least I wasn't then.

“And then in step three, rather than fine tuning on a reasonably specific task classification, let's fine tune on a, on a RLHF task classification. And so that was really, that was really key, you know, so I was kind of like out of the NLP field for a few years there because yeah, it just felt like, I don't know, pushing uphill against this vast tide, which I was convinced was not the right direction, but who's going to listen to me, you know, cause I, as you said, I don't have a PhD, not at a university, or at least I wasn't then.”
Speaker
Jeremy Howard
Publisher
Latent Space
Search evidence