High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Reiner Pope: commitment

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“I can motivate the batch at least a little bit. We will see exactly why batch is such a favorable optimization.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
commitment
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…I’ll just interrupt from time to time to ask some very naive questions or to clarify some basic points. For the audience, you’re not serving one user at a time. The batch refers to the fact that you’re serving many different users at the same time, and that’s a whole batch. I can motivate the batch at least a little bit. We will see exactly why batch is such a favorable optimization. What will turn out to be the case is that if you do not batch together many users, the cost and the economics you get can be a thousand times worse than if you do batch many users together. We’ll be able to see that quite explicitly. Then, number of active parameters. If I look at, for example, a DeepSeek model, the DeepSeek V3 model has about 37 billion active parameters, and 700 billion total parameters. We’re focusing on just the ones that are active for a single AI token. We’re modeling compute performance. I’m going to keep writing equals, but in all of these cases, you can think of this time as being at least this much, and maybe there will be some terms we ignored. On the memory side, what do we need to do with memory? We need to fetch all of the weights, so there is some time to fetch the total number of parameters, not just the active parameters. There’s weight fetch time, and then in addition, there’s a KV cache fetch time. This actually depends on batch size. For every element of the batch, we have to fetch an entire context length worth of tokens, and there’s a size per token, bytes for one token. This is a model parameter. Maybe just backing up, let’s explain what the KV cache is real quick.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence