High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Mark Zuckerberg: evaluation

29 Apr 2025 Dwarkesh Podcast Mark Zuckerberg — AI will write most Meta code in 18 months

“I'm very interested in studying this because I think one of the main things that's interesting about open source is the ability to distill models.”

— Mark Zuckerberg

Source trail

Everything needed to verify it.

Speaker
Mark Zuckerberg
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2025
Publisher
Dwarkesh Podcast

Transcript context

…As in, it's important that people are building for Llama rather than for LLMs in general, because that will determine what the standard is in the future. Look, I think these models encode values and ways of thinking about the world. We had this interesting experience early on, where we took an early version of Llama and translated it. I think it was French, or some other language. The feedback we got from French people was, "This sounds like an American who learned to speak French. It doesn’t sound like a French person." And we were like, “what do you mean, does it not speak French well?” No, it speaks French fine. It was just that the way it thought about the world seemed slightly American. So I think there are these subtle things that get built into the models. Over time, as models get more sophisticated, they should be able to embody different value sets across the world. So maybe that's not a particularly sophisticated example, but I think it illustrates the point. Some of the stuff we've seen in testing some of the models, especially coming out of China, have certain values encoded in them. And it’s not just a light fine-tune to change that. Now, language models — or something that has a kind of world model embedded in it — have more values. Reasoning, I guess, you could say has values too. But one of the nice things about reasoning models is they're trained on verifiable problems. Do you need to be worried about cultural bias if your model is doing math? Probably not. I think the chance that some reasoning model built elsewhere is going to incept you by solving a math problem in a devious way seems low. But there's a whole different set of issues around coding, which is the other verifiable domain. You need to worry about waking up one day and if you're using a model that has some tie to another government, can it embed vulnerabilities in code that their intelligence organizations could exploit later? In some future version you're using a model that came from another country and it's securing your systems. Then you wake up and everything is just vulnerable in a way that that country knows about and you don’t. Or it turns on a vulnerability at some point. Those are real issues. I'm very interested in studying this because I think one of the main things that's interesting about open source is the ability to distill models. For most people, the primary value isn't just taking a model off the shelf and saying, "Okay, Meta built this version of Llama. I'm going to take it and I'm going to run it exactly in my application." No, your application isn't doing anything different if you're just running our thing. You're at least going to fine-tune it, or try to distill it into a different model. When we get to stuff like the Behemoth model, the whole value is being able to take this very high amount of intelligence and distill it down into a smaller model that you're actually going to want to run. This is the beauty of distillation. It's one of the things that I think has really emerged as a very powerful technique over the last year, since the last time we sat down. ually going to want to run. This is the beauty of distillation. It's one of the things that I think has really emerged as a very powerful technique over the last year, since the last time we sat down. I think it’s worked better than most people would have predicted. You can basically take a model that's much bigger, and capture probably 90 or 95% of its intelligence, and run it in something that's 10% of the size. Now, do you get 100% of the intelligence? No. But 95% of the intelligence at 10% of the cost is pretty good for a lot of things. The other thing that's interesting is that now, with this more varied open-source community, it's not just Llama. You have other models too. You have the ability to distill from multiple sources. So now you can basically say, "Okay, Llama’s really good at this. Maybe its architecture is really good because it's fundamentally multimodal, more inference-friendly, more efficient. But let’s say this other model is better at coding." Okay, great. You can distill from both of them and build something that's better than either individually, for your own use case. That's cool. But you do need to solve the security problem of knowing that you can distill it in a way that's safe and secure. This is something that we've been researching and have put a lot of time into. What we've basically found is that anything that's language is quite fraught. There's just a lot of values embedded into it. Unless you don't care about taking on the values from whatever model you're distilling from, you probably don't want to just distill a straight language world model. On reasoning, though, you can get a lot of the way there by limiting it to verifiable domains, and running code cleanliness and security filters. Whether it's using Llama Guard open source, or the Code Shield open source tools that we've done, things that allow you to incorporate different input into your models and make sure that both the input and the output are secure. Then it’s just a lot of red teaming. It’s having experts who are looking at the model and asking, "Alright, is this model doing anything after distillation that we don't want?" I think with the combination of those techniques, you can probably distill on the reasoning side for verifiable domains quite securely. That's something I'm pretty confident about and something we've done a lot of research around. But I think this is a very big question. How do you do good distillation? Because there’s so much value to be unlocked. But at the same time, I do think there is some fundamental bias embedded in different models.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence