Evidence receipt / evaluation
Published · transcript-backedReiner Pope: evaluation
29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served
“It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.”
Source trail
Everything needed to verify it.
- Speaker
- Reiner Pope
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 29 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…When I do a forward pass… Let me draw how the autoregressive inference works. This is during decode. If I have a bunch of text tokens… I’m drawing a tensor because ultimately the tokens are represented as a tensor in some embedding dimension. In this direction, I have the sequence length. The work of running a decode is that I have to run each token through a whole bunch of matrix multiplies over a bunch of different layers. In general, I’m going to have to do that work over all of these tokens. But one step of decode is to produce just this one additional token up here. What I’m going to do there is run a full forward pass of multiplying by all of the weight matrices in the entire model. But then I’ve got this attention mechanism where this token is looking at all of the past tokens, and what is it looking at specifically? It is looking at some internal representation that the model has produced of the tokens, and we call that the KV cache. This process of this single token attending to all of the history of tokens is attention. It is mostly dominated by memory fetches rather than matrix multiplies. So we’ve got the amount of memory that we’re fetching shown over here, and then this is of course just divided by the memory bandwidth, so the memory bytes per second. In fact, these equations here are enough for us to now draw some fit lines. The things that we’d like to look at are sensitivity to batch, and then also, which we’ll draw separately, to context length. We said that the big effect you can get is some trade-off in latency versus cost in batch size. Let’s draw them out. I think there are just really two graphs that we want to draw. We’ll first draw batch size versus time here. When we look at the shape of this, we’ve got a maximum of the sum and then another term. Let’s look at these terms one by one and how they scale: the time for compute and memory, and how they show up. Let’s first look at this compute time. This is just purely linear in batch size with no offset, so it is some curve like this. This is t compute. On the memory side, we’ve got some portion here that is just this constant in some base offset here, which is the weight fetch. Finally, we have this term here, which is the KV fetch, which is pretty linear in batch size, and so it looks like that. The sum of this plus this maxed with this… Let’s at least first draw the sum. The two memory times in conjunction end up looking on this curved slope like this. Then the overall maximum is—I’ll draw a little thicker here—the maximum of these two curves. What does this mean? This is a latency plot. If I grow my batch size, initially I get some not very strong dependence on batch size, so there is some lower bound on latency here. This already partially answers the question. For a given hardware configuration—and we can talk about varying the hardware configuration—there is a lower bound on latency. lower bound on latency here. This already partially answers the question. For a given hardware configuration—and we can talk about varying the hardware configuration—there is a lower bound on latency. It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that. It seems like the way you’ve drawn the slopes for compute time and how the KV grows—and what implication the KV has on memory time—…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.