Evidence receipt / evaluation
Published · transcript-backedJosh Albrecht: evaluation
25 Jun 2024 Latent Space State of the Art: Training >70B LLMs on 10,000 H100 clusters
“I think ours is actually quite a bit lower than that, probably because we've taken the time to like dig into a large, maybe larger number than we should have of these failures and get to the root cause of it and be like, oh, okay, like that's exactly what's going wrong.”
Source trail
Everything needed to verify it.
- Speaker
- Josh Albrecht
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Jun 2024
- Publisher
- Latent Space
Transcript context
…Something we discussed in the pre-show was that you had a rule of thumb for your cluster of reliability. You say here in the post, by and large, you expect around 3% of your machines to break every week. So you're basically going to turn through all your machines in a year. As it says in the post. So that would be true if it was a uniform failure like that. But as it says in the post, it's usually these kind of problematic nodes. And to be clear, that is the number that we've heard from other people is like they're having about 3%. I don't think we're experiencing failure rates that are that high. I think ours is actually quite a bit lower than that, probably because we've taken the time to like dig into a large, maybe larger number than we should have of these failures and get to the root cause of it and be like, oh, okay, like that's exactly what's going wrong. How do we fix this?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.