Evidence receipt / belief
Published · transcript-backedNathan Lambert: belief
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Lambert
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…Pre-training, I’m using some of the same words to really get the message across is you’re doing what is called autoregressive prediction to predict the next token in a series of documents. This is done over standard practice is trillions of tokens, so this is a ton of data that is mostly scraped from the web. And some of DeepSeek’s earlier papers, they talk about their training data being distilled for math. I shouldn’t use this word yet, but taken from Common Crawl and that’s a public access that anyone listening to this could go download data from the Common Crawl website. This is a crawler that is maintained publicly. Yes, other tech companies eventually shift to their own crawler and DeepSeek likely has done this as well as most frontier labs do. But this sort of data is something that people can get started with and you’re just predicting text in a series of documents. This can be scaled to be very efficient and there’s a lot of numbers that are thrown around in AI training like how many floating-point operations or flops are used. And then you can also look at how many hours of these GPUs that are used. And it’s largely one loss function taken to a very large amount of compute usage. You set up really efficient systems and then at the end of that you have the base model and pre-training is where there is a lot more of complexity in terms of how the process is emerging or evolving and the different types of training losses that you’ll use. I think this is a lot of techniques grounded in the natural language processing literature. The oldest technique which is still used today is something called instruction tuning or also known as supervised fine-tuning. These acronyms will be IFT or SFT. People really go back and forth throughout them, and I’ll probably do the same, which is where you add this formatting to the model where it knows to take a question that is, explain the history of the Roman Empire to me or a sort of question you’ll see on Reddit or Stack Overflow. And then the model will respond in a information-dense but presentable manner. The core of that formatting is in this instruction tuning phase. And then there’s two other categories of loss functions that are being used today. One I’ll classify as preference fine-tuning. Preference fine-tuning is a generalized term for what came out of reinforcement learning from human feedback, which is RLHF. This reinforcement learning from human feedback is credited as the technique that helped ChatGPT break through. It is a technique to make the responses that are nicely formatted like these Reddit answers more in tune with what a human would like to read. dited as the technique that helped ChatGPT break through. It is a technique to make the responses that are nicely formatted like these Reddit answers more in tune with what a human would like to read. This is done by collecting pairwise preferences from actual humans out in the world to start and now AIs are also labeling this data and we’ll get into those trade-offs. And you have this contrastive loss function between a good answer and a bad answer. And the model learns to pick up these trends. There’s different implementation ways. You have things called reward models. You could have direct alignment algorithms. There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API. This is the idea that you use the techniques of reinforcement learning, which is a whole framework of AI. There’s a deep literature here to summarize, it’s often known as trial and error learning or the subfield of AI where you’re trying to make sequential decisions in a certain potentially noisy environment. There’s a lot of ways we could go down that, but fine-tuning language models where they can generate an answer and then you check to see if the answer matches the true solution. For math or code you have an exactly correct answer for math, you can have unit tests for code. And what we’re doing is we are checking the language model’s work and we’re giving it multiple opportunities on the same questions to see if it is right. And if you keep doing this, the models can learn to improve in verifiable domains to a great extent. It works really well. It’s a newer technique in the academic literature. It’s been used at frontier labs in the US that don’t share every detail for multiple years. This is the idea of using reinforcement learning with language models and it has been taking off especially in this DeepSeek moment. And we should say that there’s a lot of exciting stuff going on again across the stack, but the post-training probably this year, there’s going to be a lot of interesting developments in the post-training. We’ll talk about it. I almost forgot to talk about the difference between DeepSeek-V3 and R1 on the user experience side. Forget the technical stuff, forget all of that, just people that don’t know anything about AI, they show up. What’s the actual experience, what’s the use case for each one when they actually type and talk to it? What is each good at and that kind of thing?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.