High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…Has that ratio changed over time as we’ve gone from model generation to model generation, where the FLOPs keep increasing? This is a hardware parameter. To what extent has the hardware changed? From A100 to H100 to B100, the FLOPs have increased substantially, the memory bandwidth has also increased substantially, and it has remained reasonably stable. We can express this one as well. This is a sparsity parameter. I might even phrase this slightly differently. Let’s solve for batch size in total. Moving this back over to the other side, we end up with batch size needs to be bigger than approximately 300 times sparsity. For example, in DeepSeek I activate 32 out of 256 experts, so this would be 8 for DeepSeek. This actually gives you a ballpark which is remarkably accurate to practice. Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say. But take this and maybe double or triple it. Okay, so it’s two to three thousand tokens per batch. But then if you included the KV cache, the implication would be that the optimal batch size...…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence