Evidence receipt / prediction
Published · transcript-backedMark Zuckerberg: prediction
29 Apr 2025 Dwarkesh Podcast Mark Zuckerberg — AI will write most Meta code in 18 months
“In general, the prediction that this would be the year open source generally overtakes closed source as the most used models out there, I think that's generally on track to be true.”
Source trail
Everything needed to verify it.
- Speaker
- Mark Zuckerberg
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 29 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…I'm interested to hear more about it. There's this impression that the gap between the best closed-source and the best open-source models has increased over the last year. I know the full family of Llama 4 models isn't out yet, but Llama 4 Maverick is at #35 on Chatbot Arena. On a bunch of major benchmarks, it seems like o4-mini or Gemini 2.5 Flash are beating Maverick, which is in the same class. What do you make of that impression? There are a few things. First, I actually think this has been a very good year for open source overall. If you go back to where we were last year, Llama was the only real, super-innovative open-source model. Now you have a bunch of them in the field. In general, the prediction that this would be the year open source generally overtakes closed source as the most used models out there, I think that's generally on track to be true. One interesting surprise — positive in some ways, negative in others, but overall good — is that it’s not just Llama. There are a lot of good ones out there. I think that's quite good. Then there's the reasoning phenomenon, which you're alluding to talking about o3, o4, and other models. There's a specialization happening. If you want a model that’s the best at math problems, coding, or different things like those tasks, then reasoning models that consume more test-time or inference-time compute in order to provide more intelligence are a really compelling paradigm. And we're building a Llama 4 reasoning model too. It'll come out at some point. But for a lot of the applications we care about, latency and good intelligence per cost are much more important product attributes. If you're primarily designing for a consumer product, people don't want to wait half a minute to get an answer. If you can give them a generally good answer in half a second, that's a great tradeoff. I think both of these are going to end up being important directions. I’m optimistic about integrating reasoning models with the core language models over time. That's the direction Google has gone in with some of the more recent Gemini models. I think that's really promising. But I think there’s just going to be a bunch of different stuff that goes on. You also mentioned the whole Chatbot Arena thing, which I think is interesting and points to the challenge around how you do benchmarking. How do you know what models are good for which things? One of the things we've generally tried to do over the last year is anchor more of our models in our Meta AI product north star use cases. The issue with open source benchmarks, and any given thing like the LM Arena stuff, is that they’re often skewed toward a very specific set of uses cases, which are often not actually what any normal person does in your product. The portfolio of things they’re trying to measure is often different from what people care about in any given product. Because of that, we’ve found that trying to optimize too much for that kind of stuff has led us astray. It’s actually not led towards the highest quality product, the most usage, and best feedback within Meta AI as people use our stuff. So we're trying to anchor our north star on the product value that people report to us, what they say that they want, and what their revealed preferences are, and using the experiences that we have. Sometimes these benchmarks just don't quite line up. I think a lot of them are quite easily gameable. y that they want, and what their revealed preferences are, and using the experiences that we have. Sometimes these benchmarks just don't quite line up. I think a lot of them are quite easily gameable. On the Arena you'll see stuff like Sonnet 3.7, which is a great model, and it's not near the top. It was relatively easy for our team to tune a version of Llama 4 Maverick that could be way at the top. But the version we released, the pure model, actually has no tuning for that at all, so it's further down. So you just need to be careful with some of these benchmarks. We're going to index primarily on the products. Dwarkesh Patel Do you feel like there is some benchmark which captures what you see as a north star of value to the user which can be be objectively measured between different models and where you'd say, "I need Llama 4 to come out on top on this”? Our benchmark is basically user value in Meta AI.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.