Evidence receipt / evaluation
Published · transcript-backedReiner Pope: evaluation
29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served
“If I think about what that means, it means I actually have time to read the HBM twice.”
Source trail
Everything needed to verify it.
- Speaker
- Reiner Pope
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 29 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…Let me make sure I understand what this is saying. I understand the unit analysis. What it’s saying is we can evacuate and replace the HBM in this amount of time. So we don’t want to be in a situation where the HBM is not big enough that we’re not actually able to write everything we want to it or take everything out of it. Or we don’t want to be in a situation where our ability to write back and forth is so small compared... There are sort of two scenarios. Why don’t we pick a latency that is bigger than 15 milliseconds? If I think about what that means, it means I actually have time to read the HBM twice. By the way, most HBM accesses are reads, not writes. It’s almost all reads because the weight matrices are read-only, and almost all of the KV cache accesses are reads. In around 30 milliseconds, I can read all of HBM twice, but what’s the point of that? I don’t want to read the weight matrices twice. I don’t want to read the KVs twice. Makes a ton of sense. A couple of quick questions. If it is the case that the optimal batch size is something like 2,000, it’s totally dependent on the sparsity, not dependent on the model size or anything.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.