High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Quentin Anthony: belief

16 Aug 2023 Latent Space The Mathematics of Training LLMs — with Quentin Anthony of Eleuther AI

“I would say that a lot of major GPT-based models use this scheme. A lot of them now are sort of going with just a pure zero scheme.”

— Quentin Anthony

Source trail

Everything needed to verify it.

Speaker
Quentin Anthony
Attribution
Verified speaker
Claim type
belief
Recorded
16 Aug 2023
Publisher
Latent Space

Transcript context

…Like the sharded optimizers plus the 3D parallelism, bringing the two things together and having this kind of mesh strategy. I would say that a lot of major GPT-based models use this scheme. A lot of them now are sort of going with just a pure zero scheme. So just a pure sharded. You just shard everything. And then since that's so easy, everyone gets an equal slice. There's no such thing as a pipeline stage. There's no such thing as what tensor should go on which GPU. Instead, we shard everything equally and treat everything equally. It's a much easier problem to debug, to checkpoint, to run training on than it is with this 3D parallel scheme. I say 3D parallel gives you the most control and also the most ways to go wrong. And depending on whether you have more engineers or whether you have more GPUs, that should decide which of these you go with. It's also not too hard, right? You've basically outlined the five or six different numbers that you need to keep in your head. And it doesn't feel impossible that if you need to achieve that level of control, you've given everybody the main levers to do it with. And that's wonderful. Definitely.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence