Evidence receipt / evaluation
Published · transcript-backedNathan Lambert: evaluation
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Lambert
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…Wait, so you’re saying I can’t make a cheap copy of Llama and pretend it’s mine, but I can do this with the Chinese model? … Hell, yeah. That’s what I’m saying. And that’s why it’s like we want this whole open language model thing, he Olmo thing is to try to keep the model where everything is open with the data as close to the frontier as possible. So we’re compute constrained, we’re personnel constrained. We rely on getting insights from people like John Schulman tells us to do URL and outputs. We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software. We talked about open source software for security. Also it’s just because you build something once and can reuse it. If you go into a new company, there’s so many benefits, but if you open source a language model, you have this data sitting around, you have this training code, it’s not like that easy for someone to come and build on and improve because you need to spend a lot on compute, you need to have expertise. So until there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like America needs this and I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around, what benefits do you get from seeing the language model data? And there’s not a lot about that. We’re going to try to launch a demo soon where you can look at an OMO model and a query and see what pre-training data is similar to it, which is legally risky and complicated, but it’s like what does it mean to see the data that the AI was trained on? It’s hard to parse. It’s terabytes of files. It’s like I don’t know what I’m going to find in there, but that’s what we need to do as an ecosystem if people want open source AI to be financially useful. We didn’t really talk about Stargate. I would love to get your opinion on what the new administration, the Trump administration, everything that’s being done from the America side and supporting AI infrastructure and the efforts of the different AI companies. What do you think about Stargate? What are we supposed to think about Stargate and does Sam have the money?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.