Evidence receipt / belief
Published · transcript-backedMark Zuckerberg: belief
18 Apr 2024 Dwarkesh Podcast Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters
“I guess our prediction going in was that it was going to asymptote more, but even by the end it was still learning.”
Source trail
Everything needed to verify it.
- Speaker
- Mark Zuckerberg
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 18 Apr 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…In the material they shared with me before, it was really interesting that you trained it on more data than is compute optimal just for training. The inference is such a big deal for you guys, and also for the community, that it makes sense to just have this thing and have trillions of tokens in there. Although one of the interesting things about it, even with the 70B, is that we thought it would get more saturated. We trained it on around 15 trillion tokens. I guess our prediction going in was that it was going to asymptote more, but even by the end it was still learning. We probably could have fed it more tokens and it would have gotten somewhat better. At some point you're running a company and you need to do these meta reasoning questions. Do I want to spend our GPUs on training the 70B model further? Do we want to get on with it so we can start testing hypotheses for Llama-4? We needed to make that call and I think we got a reasonable balance for this version of the 70B. There'll be others in the future, the 70B multimodal one, that'll come over the next period. But that was fascinating that the architectures at this point can just take so much data. That's really interesting. What does this imply about future models? You mentioned that the Llama-3 8B is better than the Llama-2 70B.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.