Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
29 Apr 2025 Dwarkesh Podcast Mark Zuckerberg — AI will write most Meta code in 18 months
“DeepSeek has the MIT license, whereas I think a couple of the contingencies in the Llama license require you to say "built with Llama" on applications using it or any model that you train using Llama has to begin with the word "Llama.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 29 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…It's a real competition. You're seeing industrial policies really play out. China is bringing online more power. Because of that, the US really needs to focus on streamlining the ability to build data centers and produce energy. Otherwise, I think we’ll be at a significant disadvantage. At the same time, some of the export controls on things like chips, I think you can see how they’re clearly working in a way. There was all the conversation with DeepSeek about, "Oh, they did all these very impressive low-level optimizations." And the reality is, they did and that is impressive. But then you ask, "Why did they have to do that, when none of the American labs did it?" It’s because they’re using partially nerfed chips that are the only ones NVIDIA is allowed to sell in China because of the export controls. DeepSeek basically had to spend a bunch of their calories and time doing low-level infrastructure optimizations that the American labs didn’t have to do. Now, they produced a good result on text. DeepSeek is text-only. The infrastructure is impressive. The text result is impressive. But every new major model that comes out now is multimodal. It's image, it's voice. Theirs isn't. Now the question is, why is that the case? I don’t think it’s because they’re not capable of doing it. It's because they had to spend their calories on doing these infrastructure optimizations to overcome the fact that there were these export controls. But when you compare Llama 4 with DeepSeek —I mean our reasoning model isn’t out yet, so the R1 comparison isn’t clear yet— but we’re basically in the same ballpark on all the text stuff that DeepSeek is doing but with a smaller model. So the cost-per-intelligence is lower with what we’re doing for Llama on text. On the multimodal side we’re effectively leading at and it just doesn’t exist in their models. So the Llama 4 models, when you compare them to what DeepSeek is doing, are good. I think people will generally prefer to use the Llama 4 models. But there’s this interesting contour where it’s clearly a good team doing stuff over there. And you're right to ask about the accessibility of power, the accessibility of compute and chips, because the work that you're seeing different labs do and the way it's playing out is somewhat downstream of that. So Sam Altman recently tweeted that OpenAI is going to release an open-source SOTA reasoning model. I think part of the tweet was that they won’t do anything silly, like say you can only use it if you have less than 700 million users. DeepSeek has the MIT license, whereas I think a couple of the contingencies in the Llama license require you to say "built with Llama" on applications using it or any model that you train using Llama has to begin with the word "Llama. " What do you think about the license? Should it be less onerous for developers? Look, we basically pioneered the open-source LLM thing. So I don't consider the license to be onerous. When we were starting to push on open source, there was this big debate in the industry. Is this even a reasonable thing to do? Can you do something that is safe and trustworthy with open source? Will open source ever be able to be competitive enough that anyone will even care? Basically, when we were answering those questions a lot of the hard work was done by the teams at Meta. There were other folks in the industry but really, the Llama models were the ones that broke open this whole open-source AI thing in a huge way. If we’re going to put all this energy into it, then at a minimum, if you're going to have these large cloud companies — like Microsoft and Amazon and Google — turn around and sell our model, then we should at least be able to have a conversation with them before they do that around what kind of business arrangement we should have. Our goal with the license, we're generally not trying to stop people from using the model. We just think that if you're one of those companies, or if you're Apple, just come talk to us about what you want to do. Let's find a productive way to do it together. I think that’s generally been fine. Now, if the whole open-source part of the industry evolves in a direction where there are a lot of other great options and the license ends up being a reason why people don’t want to use Llama, then we’ll have to reevaluate the strategy. What it makes sense to do at that point. But I don’t think we’re there. That’s not, in practice, something we’ve seen, companies coming to us and saying, “We don’t want to use this because your license says if you reach 700 million people, you have to come talk to us.” So far, that’s been more something we’ve heard from open-source purists like, “Is this as clean of an open-source model as you’d like it to be?” That debate has existed since the beginning of open source. All the GPL license stuff versus other things, do you need to make it so that anything that touches open source has to be open source too? Or can people take it and use it in different ways? I'm sure there will continue to be debates around this. But if you’re spending many billions of dollars training these models, I think asking the other companies — the huge ones that are similar in size and can easily afford to have a relationship with us — to talk to us before they use it seems like a pretty reasonable thing.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.