Evidence receipt / belief
Published · transcript-backedMark Zuckerberg: belief
18 Apr 2024 Dwarkesh Podcast Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters
“I think you'll get smaller versions. One thing is that I think 8B isn’t quite small enough for a bunch of use cases.”
Source trail
Everything needed to verify it.
- Speaker
- Mark Zuckerberg
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 18 Apr 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…What is the community fine tune of Llama-3 that you're most excited for? Maybe not the one that will be most useful to you, but the one you'll just enjoy playing with the most. They fine-tune it on antiquity and you'll just be talking to Virgil or something. What are you excited about? I think the nature of the stuff is that you get surprised. Any specific thing that I thought would be valuable, we'd probably be building. I think you'll get distilled versions. I think you'll get smaller versions. One thing is that I think 8B isn’t quite small enough for a bunch of use cases. Over time I'd love to get a 1-2B parameter model, or even a 500M parameter model and see what you can do with that. If with 8B parameters we’re nearly as powerful as the largest Llama-2 model, then with a billion parameters you should be able to do something that's interesting, and faster. It’d be good for classification, or a lot of basic things that people do before understanding the intent of a user query and feeding it to the most powerful model to hone in on what the prompt should be. I think that's one thing that maybe the community can help fill in. We're also thinking about getting around to distilling some of these ourselves but right now the GPUs are pegged training the 405B. So you have all these GPUs. I think you said 350,000 by the end of the year.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.