Evidence receipt / evaluation
Published · transcript-backedReiner Pope: evaluation
29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served
“A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.”
Source trail
Everything needed to verify it.
- Speaker
- Reiner Pope
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 29 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…Makes sense. What is the cost of these different ways of resynthesizing the KV cache? Computing it from scratch is based on my GPU time. I have to do a certain amount of multiplies, of GPU time that I spend in order to produce it. Storing in HBM. This really goes as my bytes per token. I need to just have some number of bytes per token, and then I need to store this in the HBM. It’s going to use up some of my HBM capacity. A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU. How do I price that? Maybe I say that the cost of it is proportional to the fraction of the HBM I’m using. There’s also times GPU dollars. Let’s just do one more memory tier and say store in DDR instead. The same kind of thing goes up for flash and for DDR. I put these in the wrong columns. I meant to make two columns. The distinction I want to make is that there is the cost to retrieve, and then there’s a cost to hold on. This is a cost per second, whereas this is an instantaneous cost. Rematerialization has a cost to retrieve and has zero cost to store it because we’ve deleted it. This is the one that I put in the wrong location. This is actually the cost just to hold on, so I will rewrite it. If we’re just storing it in HBM, it has this sort of cost profile. If we store in DDR, it’s actually going to take some time. We get the same thing here: bytes per token over DDR capacity times DDR cost per second. But now this has a cost to retrieve that is higher than the HBM because we need to copy it into the HBM. So this is bytes per token over DDR bandwidth. And then this consumes some amount of the DDR as well. And every scale-up has DDR and flash?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.