Evidence receipt / belief
Published · transcript-backedNathan Lambert: belief
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“Even things that if you’re an expert, things that are close to the fringe of knowledge, they will still be fairly good at, I think.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Lambert
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…And we should say that there’s a lot of exciting stuff going on again across the stack, but the post-training probably this year, there’s going to be a lot of interesting developments in the post-training. We’ll talk about it. I almost forgot to talk about the difference between DeepSeek-V3 and R1 on the user experience side. Forget the technical stuff, forget all of that, just people that don’t know anything about AI, they show up. What’s the actual experience, what’s the use case for each one when they actually type and talk to it? What is each good at and that kind of thing? Let’s start with DeepSeek-V3, again it’s more people would tried something like it. You ask it a question, it’ll start generating tokens very fast and those tokens will look like a very human legible answer. It’ll be some sort of markdown list. It might have formatting to help you draw to the core details in the answer and it’ll generate tens to hundreds of tokens. A token is normally a word for common words or a sub word part in a longer word, and it’ll look like a very high quality Reddit or Stack Overflow answer. These models are really getting good at doing these across a wide variety of domains, I think. Even things that if you’re an expert, things that are close to the fringe of knowledge, they will still be fairly good at, I think. Cutting edge AI topics that I do research on, these models are capable for study aid and they’re regularly updated. Where this changes is with the DeepSeek- R1, what is called these reasoning models is when you see tokens coming from these models to start, it will be a large chain of thought process. We’ll get back to chain of thought in a second, which looks like a lot of tokens where the model is explaining the problem. The model will often break down the problem and be like, okay, they asked me for this. Let’s break down the problem. I’m going to need to do this. And you’ll see all of this generating from the model. It’ll come very fast in most user experiences. These APIs are very fast, so you’ll see a lot of tokens, a lot of words show up really fast, it’ll keep flowing on the screen and this is all the reasoning process. And then eventually the model will change its tone in R1 and it’ll write the answer where it summarizes its reasoning process and writes a similar answer to the first types of model. But in DeepSeek’s case, which is part of why this was so popular even outside the AI community, is that you can see how the language model is breaking down problems. And then you get this answer, on a technical side they train the model to do this specifically where they have a section which is reasoning, and then it generates a special token, which is probably hidden from the user most of the time, which says, okay, I’m starting the answer. The model is trained to do this two stage process on its own. If you use a similar model in say, OpenAI, OpenAI’s user interface is trying to summarize this process for you nicely by showing the sections that the model is doing and it’ll click through, it’ll say breaking down the problem, making X calculation, cleaning the result, and then the answer will come for something like OpenAI. Maybe it’s useful here to go through an example of a DeepSeek-R1 reasoning.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.