Evidence receipt / belief
Published · transcript-backedQuentin Anthony: belief
16 Aug 2023 Latent Space The Mathematics of Training LLMs — with Quentin Anthony of Eleuther AI
“I would say that a lot of major GPT-based models use this scheme. A lot of them now are sort of going with just a pure zero scheme.”
Source trail
Everything needed to verify it.
- Speaker
- Quentin Anthony
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 16 Aug 2023
- Publisher
- Latent Space
Transcript context
…Like the sharded optimizers plus the 3D parallelism, bringing the two things together and having this kind of mesh strategy. I would say that a lot of major GPT-based models use this scheme. A lot of them now are sort of going with just a pure zero scheme. So just a pure sharded. You just shard everything. And then since that's so easy, everyone gets an equal slice. There's no such thing as a pipeline stage. There's no such thing as what tensor should go on which GPU. Instead, we shard everything equally and treat everything equally. It's a much easier problem to debug, to checkpoint, to run training on than it is with this 3D parallel scheme. I say 3D parallel gives you the most control and also the most ways to go wrong. And depending on whether you have more engineers or whether you have more GPUs, that should decide which of these you go with. It's also not too hard, right? You've basically outlined the five or six different numbers that you need to keep in your head. And it doesn't feel impossible that if you need to achieve that level of control, you've given everybody the main levers to do it with. And that's wonderful. Definitely.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.